AI-Assisted Type B / Type C Meeting Q&A Preparation
It is forty days before an End-of-Phase-2 meeting with the FDA Office of New Drugs, and the sponsor's regulatory lead is staring at a list of eleven proposed questions and a briefing book that took six weeks to assemble. The briefing book states the sponsor's positions. The Q&A binder, the document that actually decides whether the meeting goes well, states what the sponsor will say when a reviewer pushes back on those positions, and right now it is a blank file. The team has asked the enterprise large language model to "generate the likely FDA questions and our answers," and eleven seconds later a confident, well-organized binder appears with thirty questions and thirty crisp answers. It reads like the work of a seasoned regulatory strategist. It is also dangerously incomplete, because it generated the questions a polite reviewer would ask, not the questions a skeptical statistical reviewer will actually ask about the multiplicity adjustment on the key secondary endpoint. This lesson is about how to use AI to build a Q&A binder that survives the meeting, by turning the model into a structured adversary rather than a flattering scribe, and by anchoring every answer to a source the sponsor can defend at the table.
What the Q&A Binder Actually Is, and Why It Decides the Meeting
The briefing book filed in Module 1.6 is the sponsor's opening statement; the Q&A binder is the sponsor's cross-examination prep, and the meeting is won or lost on the second of those two documents. An FDA Type B End-of-Phase-2 meeting or a Type C statistical-issues meeting is not a presentation, it is a structured discussion in which the agency's reviewers probe the positions the sponsor has staked out, and the sponsor's team has minutes, sometimes seconds, to give an answer that is accurate, consistent with the briefing book, and consistent with what every other team member would say. The Q&A binder is the artifact that makes those answers fast, aligned, and pre-verified. A good binder anticipates the question, states the sponsor's position, gives the supporting evidence with a precise source citation, and flags the fallback position if the agency does not accept the primary one. A weak binder lists the easy questions and gives answers that sound reasonable in the room and fall apart when the reviewer asks the obvious follow-up.
The reason AI is genuinely useful here is that the binder is a high-volume, pattern-rich document, and the model has seen thousands of meeting preparations in its training corpus. It can produce the structural scaffolding of a Q&A entry instantly, it can cluster the sponsor's positions into question themes, and most valuably it can role-play the reviewer who will challenge each position. The reason AI is dangerous here is the same reason it is dangerous in a Module 2.5 efficacy summary: it completes the pattern of a plausible question and a plausible answer, and a plausible answer that cites a subgroup result the sponsor never actually pre-specified is indistinguishable, on the page, from a defensible one. The discipline of this lesson is to use the model's pattern fluency to widen the question set and stress-test the positions, while never letting an unverified factual claim or an uncited cross-reference into an answer the sponsor will speak aloud to a reviewer.
The Meeting Type Determines the Reviewer Personas
Before the model drafts a single question, it has to know which meeting it is preparing for, because the meeting type determines who is on the agency side of the table and therefore which personas should stress-test the binder. An End-of-Phase-2 meeting, formally a Type B meeting scheduled within seventy days of the meeting request, is a broad discussion of the Phase 3 program design, the proposed pivotal endpoints, the statistical analysis plan at a strategic level, and the adequacy of the safety database. The reviewers in the room typically include the FDA Office of New Drugs clinical reviewer who owns the indication, a statistical reviewer from the Office of Biostatistics, and frequently a clinical pharmacology reviewer; for a product with a meaningful manufacturing question, an Office of Pharmaceutical Quality reviewer may attend or submit written comments. A Type C meeting, scheduled within seventy-five days of the request, is usually narrower and issue-specific, and a Type C statistical-issues meeting may be dominated entirely by the statistical reviewer probing the multiplicity strategy, the estimand framework under ICH E9(R1), the handling of intercurrent events, and the proposed alpha allocation.
This matters because the single most common failure of an AI-generated Q&A binder is that it produces the generalist's questions and misses the specialist's questions. A model asked for "likely FDA questions" will generate the clinical reviewer's framing, because that is the most statistically common framing in the training corpus, and it will under-represent the statistical reviewer's framing, which is more technical and less abundantly documented. The fix is to make the persona explicit. You instruct the model to answer as a named persona with a defined mandate: the FDA Office of New Drugs clinical reviewer focused on clinical meaningfulness and the totality of the benefit-risk picture, the Office of Biostatistics reviewer focused on Type I error control and estimand alignment, and the Office of Pharmaceutical Quality reviewer focused on whether the proposed commercial process and the clinical material are linked. Each persona generates a different and largely non-overlapping set of questions, and the union of those sets is the binder you actually need.
Building the Reviewer-Persona Stress-Test
The core technique of this lesson is the reviewer-persona stress-test, in which the model is prompted not to help the sponsor but to attack the sponsor's position from inside a specific reviewer's worldview. The mechanics matter. You give the model the sponsor's proposed position, the supporting evidence the sponsor intends to cite, and a system prompt that pins it to a persona with an explicit adversarial mandate: "You are an FDA Office of Biostatistics reviewer. You are skeptical of multiplicity strategies that protect the primary endpoint but leave key secondary endpoints unprotected. For the following sponsor position, generate the three hardest questions you would ask, the assumption in the sponsor's position you would most want to test, and the answer that would make you escalate to a non-agreement in the meeting minutes." The model, freed from the instinct to be helpful, produces questions the sponsor's own team is too close to the program to ask.
The output of the stress-test is not the binder; it is the raw material for the binder. The persona will generate questions that are sharp and questions that are off-target, because it is pattern-completing a skeptical reviewer, not reasoning from the actual regulatory precedent for this indication. The sponsor's team triages the generated questions into three buckets: questions the binder must answer because they are both likely and hard, questions worth a fallback position because they are plausible, and questions to discard because they misread the program. The value is asymmetric and worth stating plainly: even if only half the generated questions survive triage, the model has surfaced challenges the team would have discovered for the first time in the meeting room, which is the most expensive place to discover them. The stress-test converts the model's tendency to invent into a feature, because in this one use the goal is to generate a wide space of possible challenges, and verification happens at the answer stage, not the question stage.
Anchoring Every Answer to a Named Source
The question side of the binder tolerates the model's invention because a wrong question costs nothing but a triage decision. The answer side tolerates none of it, because an answer is a commitment the sponsor speaks aloud to a reviewer, and an answer built on a fabricated source is a credibility failure in real time at the table. Every answer in the binder must resolve to a named, verifiable source: a specific table in the briefing book, a specific result in a Clinical Study Report aligned to ICH E3, a specific section of the statistical analysis plan, a prior agreement recorded in the minutes of an earlier meeting, or a specific provision of the relevant FDA guidance. The model, left unconstrained, will write an answer that says "as demonstrated in our Phase 2b dose-response analysis, the selected dose shows the optimal benefit-risk profile," and that sentence may be cleanly true, may overstate a trend the data only suggested, or may reference a dose-response analysis that was exploratory and not powered, and the three are indistinguishable in the model's even, confident prose.
The control is the same source-link discipline taught for clinical writing, applied with even less tolerance because the audience is live. You require the model to attach a source citation to every factual claim in every answer, and you reject any answer sentence that does not carry one. Then a human reconciles each citation against the actual source before the answer enters the binder, exactly as a reviewer at the table will mentally reconcile the answer against the briefing book in front of them. A particular trap in meeting prep is the cross-meeting consistency claim: an answer that says "consistent with the agreement reached at our pre-IND meeting" is asserting the content of a prior interaction, and the model will produce that phrase whether or not the agreement exists, because it is a statistically natural thing to say in a regulatory answer. The sponsor who speaks a fabricated prior agreement to the FDA has done something far worse than cite a nonexistent table; they have misrepresented the regulatory history, and the reviewer who corrects them in the room has just learned to distrust every other answer in the binder.
The Multiplicity Question the Model Will Miss
Walk through the specific failure that opened this lesson, because it is the canonical example of where the generalist model falls short and the specialist persona earns its place. The sponsor's Phase 3 design has a primary endpoint of progression-free survival and three key secondary endpoints, of which overall survival is the one the agency and the eventual label will care about most. The sponsor's statistical analysis plan protects the primary endpoint with a clean alpha allocation but handles the key secondaries with a hierarchical testing procedure whose ordering puts overall survival third, behind two endpoints that are more likely to reach significance. A generalist model asked for likely FDA questions will ask whether the primary endpoint is clinically meaningful and whether the safety database is adequate, both reasonable and both already in the team's mind. It will very likely not ask the question the Office of Biostatistics reviewer will ask first: why is overall survival, the endpoint of greatest regulatory interest, positioned where it is least likely to be reached with protected alpha, and what is the sponsor's justification under the estimand framework of ICH E9(R1) for that ordering.
The statistical-reviewer persona, prompted adversarially, finds this immediately, because finding the weak link in a multiplicity strategy is the core of that reviewer's job and the corpus of statistical review is full of exactly this challenge. Once the question is surfaced, the binder can do its real work: it states the sponsor's position on the testing hierarchy, cites the specific section of the statistical analysis plan that defines it, gives the clinical and operational rationale for the ordering, and prepares the fallback position the sponsor will offer if the agency signals that the ordering is unacceptable. That fallback, the willingness to re-order the hierarchy or to discuss an alternative alpha-allocation scheme, is the difference between a meeting that ends in a documented agreement and a meeting that ends with the agency asking the sponsor to come back after redesigning the analysis. The model did not solve the statistical problem; the model surfaced the question early enough that the sponsor's biostatistician could solve it before the meeting rather than during it.
Documenting the Binder for the Audit Trail and the Minutes
The Q&A binder is a working document, but the AI involvement in producing it still falls under the same documentation discipline as any AI-assisted regulatory artifact, and there is a specific reason the audit trail matters here beyond general good practice. The meeting produces official FDA minutes, and those minutes record the agreements and non-agreements that will govern the rest of the program; a position the sponsor stated in the meeting, sourced from an AI-generated answer that turned out to be wrong, becomes part of the regulatory record and may have to be corrected in writing afterward, which is a costly and credibility-damaging exercise. For that reason, the binder's answers should carry not only their source citations but a record of human verification: who confirmed each answer against its source, and when. The AI-generation step itself should be logged with the prompt, the persona system prompts used for the stress-test, the model and version, and the sources loaded, consistent with the transparency expectation in the FDA-EMA Guiding Principles of Good AI Practice and the traceability expectation under 21 CFR Part 11.
There is also a practical reason rooted in the meeting workflow. Type B and Type C meeting requests are accompanied by a briefing package due no later than thirty days before the meeting, and in the period between submitting the package and holding the meeting the sponsor often receives the FDA's preliminary written responses to the proposed questions. Those preliminary responses reshape the meeting, because the discussion focuses on the questions where the agency's written position and the sponsor's position diverge. A well-built Q&A workflow uses the model a second time at this point: feed the FDA preliminary responses back in, and ask the personas to generate the follow-up questions the agency is now likely to press in the live discussion, given what they have already said in writing. This second pass is where the binder earns its keep, because it prepares the sponsor for the specific live discussion rather than the generic one, and it is only possible if the first-pass binder was built as a structured, sourced, version-controlled artifact rather than a one-time generated document that no one can update.
What This Means for the Regulatory Lead on Monday
The regulatory lead who opened a blank Q&A file is not wrong to reach for the model; they are wrong only if they accept what it first produces. The workflow that holds up has four named moves. First, define the meeting type and its reviewer personas explicitly, so the binder is stress-tested by the specialists who will actually be in the room, not by a generalist average of all reviewers. Second, run the reviewer-persona stress-test to widen the question set far beyond what the team would generate alone, and triage the output into must-answer, fallback-worthy, and discard, treating the model's tendency to invent questions as the asset it is in this one context. Third, anchor every answer to a named, verified source, reject any answer sentence without a citation, and treat cross-meeting consistency claims as the highest-risk assertions in the binder because a fabricated prior agreement is spoken aloud to the agency. Fourth, log the AI involvement and the human verification, and run the model a second time against the FDA preliminary responses to prepare for the live discussion that those responses will shape.
The deeper point is that meeting preparation is the one regulatory writing task where the model's greatest weakness, its willingness to generate plausible content beyond the evidence, is also its greatest strength when pointed in the right direction. Pointed at the answers, that willingness fabricates commitments the sponsor cannot keep. Pointed at the questions, that same willingness surfaces the challenges the sponsor's own expertise has blinded them to. The skill is knowing which side of the binder tolerates invention and which side tolerates none, and building the workflow so the model invents only where invention is cheap and verification catches everything where it is not. The binder the sponsor carries into the room should read as if a skeptical reviewer helped write it, because in the most useful sense one did.
Key Takeaways
- The Q&A binder, not the briefing book, decides the meeting, and it is the document AI is both most useful and most dangerous for. The briefing book states the sponsor's positions; the binder states what the sponsor will say when a reviewer challenges them, with each answer anchored to a defensible source.
- The meeting type determines the reviewer personas, and a generalist prompt misses the specialist's questions. An End-of-Phase-2 Type B meeting brings the clinical, statistical, and sometimes OPQ reviewers; a Type C statistical-issues meeting may be dominated by the Office of Biostatistics reviewer probing multiplicity and estimands under ICH E9(R1).
- The reviewer-persona stress-test turns the model into a structured adversary, and this is the one context where its tendency to invent is an asset. Prompt each persona to attack the sponsor's position; triage the generated questions into must-answer, fallback-worthy, and discard, because surfacing a hard question before the meeting is far cheaper than discovering it in the room.
- The answer side of the binder tolerates no invention, and cross-meeting consistency claims are the highest-risk assertions. Require a source citation on every factual claim, reject any uncited answer sentence, and never let the model assert a prior agreement the sponsor cannot verify, because a fabricated prior agreement is spoken aloud to the FDA.
- Log the AI involvement and run a second pass against the FDA preliminary responses. Capture prompts, persona system prompts, model and version, sources, and human verification per the FDA-EMA Guiding Principles and 21 CFR Part 11, then re-stress-test against the agency's written responses to prepare for the specific live discussion they will shape.
Skill.re