AI-Assisted FDA Form Drafting: 1571, 1572, 3674, 356h
The submission publisher opens the package at 4:40 on a Friday, two business days before the planned IND filing, and pulls up Form 1572 for the lead investigator at a 31-site Phase 2 study. Section 6 lists three sub-investigators. The site's delegation log, which the publisher cannot see from inside the eCTD, lists nine people with delegated study duties, six of whom meet the FDA definition of a sub-investigator. The AI that pre-populated the form did exactly what it was asked: it filled Form 1572 from the structured data in Veeva Vault RIM, and Vault RIM held three names because that is what the site coordinator entered into the system three months ago. The form is internally consistent, perfectly formatted, and wrong. It is wrong in a way that no spell-check, no PDF validator, and no eCTD technical-validation pass will ever catch, because the gap is not inside the form. The gap is between the form and a source the form was never reconciled against. This lesson is about the precise place AI helps in FDA form drafting, the precise place it cannot, and how to build the verification gate that turns "the AI filled the form" into "the named regulatory author certified the form," for Forms 1571, 1572, 3674, and 356h.
What These Four Forms Actually Are, and Why AI Touches Them at All
Four FDA forms carry the administrative spine of the two submissions that matter most to a sponsor's regulatory operations team: the IND and the marketing application. Form FDA 1571 is the Investigational New Drug Application cover form, the master index and transmittal that fronts every IND and every subsequent IND submission, where box 11 declares the contents of the submission and box 15 states the purpose of this particular sequence. Form FDA 1572 is the Statement of Investigator, the binding commitment an individual principal investigator signs under 21 CFR 312.53(c), naming the investigator, the site, the IRB, the sub-investigators, and the protocols covered. Form FDA 3674 is the Certification of Compliance with ClinicalTrials.gov registration requirements under Section 801 of the FDA Amendments Act, the box-checked attestation that the applicable clinical trials in the submission have been registered. Form FDA 356h is the Application to Market a New or Abbreviated New Drug or Biologic, the cover application for an NDA, BLA, or ANDA, with its long checklist of which CTD modules and components are present.
These forms are structured, repetitive, metadata-driven, and high-volume. A single large Phase 3 program may generate dozens of 1572s across its sites and refresh them every time an investigator, a sub-investigator, or an IRB changes. That profile is exactly where AI assistance pays. The fields are predictable, the source data already lives in a Regulatory Information Management system, and the cost of a transcription error is high enough to justify a machine that pre-populates and cross-checks faster than a human reading line by line. But the same profile that makes the forms ideal for automation is what makes the failure mode invisible: a form that is structurally flawless feels finished, and the reviewer who built it stops looking precisely when the most dangerous error, a data inconsistency against an unloaded source, is the only kind left.
The Three Jobs: Pre-Populate, Validate Against Current Form Versions, Flag Before the Publisher
AI assistance on these forms decomposes into three distinct jobs, and keeping them distinct is the discipline. The first job is pre-population: pulling investigator names, site addresses, IRB details, protocol numbers, NCT numbers, application numbers, and submission-purpose text out of Veeva Vault RIM metadata and placing each value in the correct field of the correct form. This is the time saver, and it is genuinely safe when, and only when, the source field in RIM is itself correct and current. The model is transcribing structured data, not generating it, which is the lowest-risk mode an LLM operates in. The risk that remains is a field-mapping error, where a value that is correct in RIM lands in the wrong box on the form, and a model that has been shown the current form's field layout will make that error rarely. The deeper risk is upstream: garbage in RIM becomes garbage on the form, rendered so cleanly that it reads as verified.
The second job is validation against the current FDA form version. FDA reissues these forms, and an expired or superseded form edition is a technical-validation finding that can cost a submission cycle. Form 1571, 1572, 3674, and 356h each carry an OMB control number and an expiration date in the form footer, and FDA periodically revises field structure, certification language, and the ClinicalTrials.gov attestation wording. The job here is to confirm that the form template in the package matches the current edition posted at the FDA Forms catalog, that no field has been added or removed since the template was cached in RIM, and that the certification language on 3674 and the application-type checkboxes on 356h reflect the current version. AI can do a structural diff between the populated form and a known-current reference, and it can flag a footer expiration date that has passed, but the authoritative source is the FDA Forms page, and the human owns the confirmation that the reference itself is current.
The third job, and the one that earns the lesson, is flagging missing or inconsistent data before the publisher gets the package. This is where AI moves from transcription to cross-reference checking, and where it delivers the most value if it is scoped correctly. A well-built check compares the protocol number on the 1572 against the protocol number in the IND's box-11 contents on the 1571, the NCT number on the 3674 against the NCT registered for that protocol, the investigator name on the 1572 against the investigator named in the protocol's site list, and the application number across all four forms. Every one of these is a consistency check the model can run in seconds, and every inconsistency it surfaces is a finding the publisher would otherwise discover late or not at all. The point of moving the check before the publisher is that the publisher's pass is technical, not semantic: the publisher confirms the package assembles and validates, not that the protocol number on the 1572 is the right protocol number.
Form 1572 Sub-Investigator Validation: The Delegation-Log Gap
Return to the Friday-afternoon 1572 with three sub-investigators listed and nine people on the delegation log. The 1572 sub-investigator field, Section 6, must list the names of the sub-investigators who will be assisting the principal investigator in the conduct of the investigation, and FDA's longstanding position, reinforced in its 2010 guidance on the 1572 and in routine bioresearch monitoring inspection findings, is that anyone delegated significant study-related duties that could affect subject safety or data integrity is a sub-investigator who belongs on the form. The recurring real-world deficiency, written up in Form 483 observations year after year, is a 1572 that omits a person who was performing protocol-specified assessments, dosing, or eligibility determinations. The omission is not caught at filing; it is caught at a for-cause or routine inspection, sometimes years later, when an inspector cross-walks the delegation log against the 1572 and finds a name on one and not the other.
This is the single most important thing to understand about AI and Form 1572: the model can validate the form against the data it is given, and the data it is given is almost never the delegation log. The delegation of authority log lives in the Trial Master File or the site's regulatory binder, often as a scanned wet-ink document, not in the RIM metadata the AI pre-populates from. So the AI populates Section 6 from the three names in RIM, validates internal consistency, finds none, and reports the form clean. The form is clean against its world and wrong against the site's. The correct workflow design makes the delegation log a required input to the 1572 check rather than an optional one: the verification gate is not "does Section 6 match RIM" but "does Section 6 reconcile to the current delegation log for this site as of this date." When the delegation log is loaded, the AI becomes genuinely useful, because cross-walking nine log entries against three form entries and flagging the six missing names is exactly the tedious, error-prone reconciliation a human does badly and a machine does well. When the log is not loaded, the AI's clean report is a false assurance, and the failure is a silence, not a flag.
There is a second 1572 trap the model handles well once scoped: the multi-protocol and multi-site investigator. An investigator running three protocols at one site needs each protocol reflected, and an investigator at one site under one 1572 cannot be silently extended to a second site without a new statement. The AI can enforce the rule that one 1572 binds one investigator at the named facility for the named protocols, and flag any attempt to reuse a 1572 across a site the investigator was not committed to, but only because that rule was written into the check. The pattern throughout is the same: the model executes a reconciliation rule reliably; the human owns whether the right rule and the right source were supplied.
Form 3674 ClinicalTrials.gov Certification Cross-Check
Form 3674 certifies compliance with the registration and results-reporting requirements of 42 U.S.C. 282(j), and the certifier is attesting, under penalty, that the applicable clinical trials referenced in the submission have been registered at ClinicalTrials.gov and that the NCT number is provided where required. The form offers three certification options: that the submission does not reference any applicable clinical trial, that it does and the requirement is met with the NCT number supplied, or that the requirement does not apply for a stated reason. The cross-check job for AI is to reconcile what the 3674 asserts against what is actually true in two external systems: the contents of the submission, which establishes whether an applicable clinical trial is referenced, and the ClinicalTrials.gov record, which establishes whether that trial is in fact registered and whether the NCT number on the form matches the registered study.
The valuable check is the three-way reconciliation. First, does the submission reference a trial that meets the "applicable clinical trial" definition, which for drugs and biologics generally means a controlled clinical investigation, other than a Phase 1 trial, of a product subject to FDA regulation. If the program includes a Phase 2 or 3 interventional trial, the 3674 cannot honestly claim that no applicable clinical trial is referenced, and an AI that has the protocol list can flag a 3674 that checked the wrong box. Second, does the NCT number on the form correspond to a real, active ClinicalTrials.gov registration for that exact protocol, not a different study of the same drug, an NCT that resolves on ClinicalTrials.gov to a study whose title, sponsor, and phase match the protocol in the submission. Third, is the registration current enough that the certification holds, since a trial that should have been registered within 21 days of first patient enrollment but was registered late, or whose results were not posted when due, creates a certification exposure the form does not show on its face.
Here the AI's reach and its limit are both sharp. It can flag the internal logical inconsistency, a 3674 claiming no applicable trial while the package contains a Phase 3 protocol, with high reliability, because both facts are inside the loaded submission. It can flag an NCT number that is malformed or absent where the certification option requires one. What it cannot do without live, verified access to ClinicalTrials.gov is confirm that the NCT actually resolves to the matching study, and a model asked to verify an NCT from memory will produce a confident answer about a registration it has not seen, which is the hallucination pattern this program returns to constantly. The defensible workflow treats the ClinicalTrials.gov match as a human-confirmed or system-integrated step, with the AI flagging candidates and a person or a validated integration confirming the resolution. The certifier signs the 3674 under penalty; the certifier, not the model, owns the attestation.
Form 1571 and 356h: The Index Forms Where Box-Level Errors Hide
Forms 1571 and 356h are different animals from 1572 and 3674: they are index and transmittal forms, and their errors are errors of contents declaration rather than of factual attestation about a person or a trial. On Form 1571, box 11 declares the contents of the submission and box 15 states the purpose of the serial submission, and the AI's job is to ensure box 11 reflects what is actually in the eCTD sequence and box 15 reflects the actual regulatory purpose of the sequence. A 1571 that says, in box 11, that the submission contains a protocol amendment, when the sequence actually carries an information amendment, is the precise classification error that the next lesson in this chapter dissects in depth, and it begins here, on the form, where the wrong box gets checked. The AI can reconcile the box-11 declaration against the eCTD section content actually present in the sequence and flag a declared content type that has no corresponding module content, or module content that is not declared.
Form 356h carries an extensive component checklist for the marketing application: which application type, which CTD modules and sections are included, whether a patent certification is present, whether a Pediatric Study Plan or assessment applies, whether the financial certifications and disclosures under 21 CFR Part 54 are attached. The high-value AI check is the completeness reconciliation, confirming that every component the 356h checklist marks as present has corresponding content in the assembled package, and that no required component for the declared application type is unmarked. A 356h that checks the box for a component the package does not contain, or that leaves unchecked a financial disclosure that the package does include, is a discrepancy a reviewer reads as carelessness, and it is exactly the kind of structured cross-reference an AI does well once the package contents are loaded as a source. The limit is the same one that runs through this entire lesson: the AI reconciles the form against the contents it can see, and the named regulatory author owns whether the contents it was given are the contents that will actually ship.
The Verification Gate and the Audit Trail That Survives an Inspection
The workflow that makes AI form drafting defensible is not the pre-population step; it is the verification gate that follows it, and the audit trail that records it. The gate has a fixed shape across all four forms. Every field that was pre-populated from RIM is reconciled to its source of truth, which for investigator and IRB data may be RIM but for sub-investigators is the delegation log, for the NCT is the ClinicalTrials.gov record, and for the box-11 contents is the actual eCTD sequence. Every cross-form consistency check the AI ran is reviewed, with each flag either resolved or explicitly dispositioned. Every form is confirmed against the current FDA form edition. And the named regulatory professional who will be accountable for the package signs the verification, because the 1572 carries an investigator's signature and the 1571 and 356h carry the sponsor's, and those signatures attach to a human who attests the form is true, not to the model that drafted it.
The audit trail records what an FDA Office of New Drugs Information Request or a bioresearch monitoring inspection will ask for: the model and version used, the system prompt and form-version reference set, the RIM and delegation-log sources loaded with their dates, the flags the AI raised, the disposition of each flag, and the human who signed the verification, all tied to the document version under change control in Veeva Vault. Under 21 CFR Part 11 and the ALCOA+ expectations this program treats as non-negotiable, the record must be attributable, contemporaneous, and complete, which means the verification cannot be reconstructed from memory after the fact; it is captured as the work is done. "The AI populated the forms from RIM" is not a record that survives an inspector cross-walking a 1572 against a delegation log. "The 1572 Section 6 was reconciled to the delegation log dated 12 May 2026, six omissions were identified and added, and the named regulatory associate verified and the investigator signed" is. The AI bought the hours; the gate and the trail are what let you keep them.
Key Takeaways
- AI assistance on Forms 1571, 1572, 3674, and 356h is three distinct jobs: pre-populate from Veeva Vault RIM metadata, validate against the current FDA form edition, and flag missing or inconsistent data before the publisher. Pre-population is the low-risk transcription mode; the cross-reference flagging is where the value and the verification both live.
- The Form 1572 sub-investigator failure is a silence, not a flag. The AI populates Section 6 from RIM, which rarely holds the delegation-of-authority log, so it reports a form clean against its world that is wrong against the site's. Make the current delegation log a required input to the 1572 check, not an optional one.
- The Form 3674 cross-check is a three-way reconciliation: applicable-trial status against the submission contents, NCT match against the ClinicalTrials.gov record, and registration currency. The AI reliably flags an internal logical inconsistency; confirming that an NCT actually resolves to the matching study is a human-confirmed or system-integrated step, never a model recollection.
- On the index forms, box-level declaration errors hide. A 1571 box-11 that declares a protocol amendment over an information amendment, or a 356h checklist that marks a component the package lacks, is reconcilable against the actual eCTD sequence, and that reconciliation is high-value AI work once the package contents are loaded as a source.
- The verification gate and the audit trail, not the pre-population, are what make the workflow defensible. Reconcile every field to its true source, disposition every flag, confirm the form edition, and capture the model, sources, flags, dispositions, and named human signature under Part 11 and ALCOA+, so the record survives a bioresearch monitoring inspection.
Skill.re