Research & Knowledge Creation
Dr. Lena Fischer is Chief AI Officer at Halden Diagnostics, a medical-imaging company with roughly 900 employees. Halden's models read chest scans and flag suspected findings for radiologists. Lena's team is good, but the company is known as a vendor, not as a contributor to the field. When she asked the board to fund a small research function, the chair pushed back with a fair question: "We buy research from others. Why should we produce it?" Lena did not have a crisp answer that day. This chapter is the answer she wishes she had given.
Why Knowledge Creation, Not Just Consumption
Most AI leaders are fluent consumers of knowledge. They read the papers, track the benchmarks, and adopt the tools, and that fluency is genuinely valuable. Creating knowledge is a different act, and at the visionary level it becomes part of the job rather than a hobby pursued by the researchers you happen to employ. When your organization publishes rigorous work, four things happen that buying research cannot buy, and each one justifies the function to a different stakeholder.
You shape the questions the field asks instead of only answering questions others posed. You attract talent who want to work where their ideas see daylight, which matters most for exactly the senior people who have the most choice about where to go. You earn a seat at the table with regulators and standards bodies who trust organizations that show their work, and that trust is difficult to acquire quickly once you need it. And you de-risk your own strategy, because the discipline of writing something down for outside scrutiny exposes the weak assumptions you would otherwise ship into production. That last benefit accrues whether or not anyone reads the paper.
For Halden specifically, the argument is sharper still. In a regulated clinical setting, the difference between a vendor and a trusted contributor is often the difference between a procurement checkbox and a design partnership with a hospital network. Those are not the same relationship and they do not carry the same margins or the same durability. Published, reproducible evidence of how Halden's models behave on edge cases is worth more to a cautious buyer than any sales deck, because it is the one form of claim the buyer can check independently.
The Forms Knowledge Creation Takes
Knowledge creation is not one thing, and treating it as "write a paper" is why many well-intentioned research programs stall. There is a spectrum of contribution, each with a different audience, effort, and payoff. A visionary leader chooses the form that fits the goal rather than defaulting to the most prestigious one, which is the most common way a research budget gets spent on the wrong artifact.
| Form | Primary audience | Effort | What it buys you |
|---|---|---|---|
| Peer-reviewed paper | Academic and clinical researchers | High | Credibility, citations, regulatory trust |
| Technical report or white paper | Practitioners and buyers | Medium | Design partnerships, sales trust |
| Open dataset or benchmark | The whole field | Medium to high | Field influence, inbound talent |
| Open-source tool | Engineers and researchers | Medium | Adoption, developer goodwill |
| Standards contribution | Regulators, standards bodies | Medium, sustained | Shaping the rules you will operate under |
| Practitioner writing (blog, talk) | Broad professional audience | Low to medium | Reach, recruiting, thought leadership |
Lena does not need Halden to produce all six. She needs to pick two or three that serve the strategy, and the choice follows from who she needs to convince. For a clinical company chasing regulatory trust, a peer-reviewed validation study plus an open benchmark of imaging edge cases does far more than a stream of blog posts, because the audiences that matter to Halden weigh peer review and reproducible artifacts heavily and weigh reach hardly at all. A developer-tools company would invert that priority entirely, and would be right to. The form follows the goal.
Deciding What Is Worth Publishing
Not everything you learn should be published. This is the judgment that separates a mature research program from a leaky one, and it is the judgment most likely to fail in the moment, because the moment usually involves an exciting result and a team that wants to share it. Some findings are genuine competitive advantage. Some touch patient data or safety-sensitive failure modes that could be misused. Some are simply not solid enough to put your name on. A visionary leader needs a decision rule so these calls are made deliberately, not in the excitement of a good result.
Use this publish-or-withhold checklist before any external release. If any answer in the withhold column is true, the default is to hold and escalate, not to ship. Making holding the default matters, because the alternative puts the burden of proof on whoever is uneasy, which in practice means the concern has to be argued under time pressure against a team that has already scheduled the announcement.
| Question | Publish if | Withhold and escalate if |
|---|---|---|
| Competitive value | It builds trust more than it hands rivals an edge | It reveals the core of a durable advantage |
| Safety and misuse | Disclosure helps defenders more than attackers | It hands a clear recipe for harm |
| Privacy | No individual is identifiable, consent covers it | Data or model outputs could re-identify people |
| Rigor | Results are reproducible and honestly caveated | You cannot reproduce it or the claim is fragile |
| Legal and contractual | Nothing bound by a partner agreement is exposed | A contract or regulation restricts disclosure |
At Halden, Lena runs every proposed publication past this checklist with legal, clinical, and communications in the room, which is deliberate: each of those functions sees a different column clearly and none of them sees all five. A benchmark of anonymized edge cases passes cleanly. A paper describing exactly how the flagging threshold is tuned per hospital gets held, because it is both competitive and a potential vector for gaming the model. The checklist does not make the decision for her, but it makes the reasoning explicit and repeatable, which means a later decision can be compared against an earlier one instead of being argued from scratch.
Doing Research With Rigor
The fastest way to destroy the credibility a research program is meant to build is to publish sloppy work. Work that does not survive scrutiny does not merely fail to build credibility, it spends the credibility you already had, because the asset you are building is precisely the assumption that your claims hold up. Rigor is not academic ceremony. It is the set of habits that make your claims survive contact with skeptical outsiders, which is the entire point of publishing. A leader does not need to run every experiment personally, but they must set and enforce the rigor bar, because under deadline pressure it is the first thing a team will quietly lower.
- Reproducibility. Someone outside your team should be able to rerun the analysis and get the same result. Publish code, data where permitted, and exact configurations.
- Honest baselines. Compare against the strongest reasonable alternative, not a weak strawman that makes your method look good.
- Report negative and null results. A finding that something did not work is real knowledge and protects others from the same dead end.
- Pre-specify your evaluation. Decide what counts as success before you look at the results, so you are not quietly moving the goalposts.
- Disclose conflicts and limitations. State who funded the work, what data it does not cover, and where the conclusions stop.
- Avoid cherry-picking. Report the full distribution of outcomes, not the single best run.
For a clinical product, rigor is also a regulatory asset rather than an overhead. When Halden validates a model, pre-specifying the evaluation protocol and reporting how the model performs on under-represented patient groups is exactly what a careful reviewer at a hospital or an agency wants to see. The same disclosure that makes a paper credible to a reviewer makes the product credible to a procurement committee. The rigor that makes research honest is the same rigor that makes it trusted.
Building a Knowledge-Creation Program
Intentions do not produce papers; systems do. A research function, even a small one, needs allocated time, a pipeline, and a gate, and it needs all three, because time without a pipeline produces unfinished work and a pipeline without a gate produces work you wish you had not shipped. Here is the structure Lena stands up at Halden, with clearly hypothetical numbers used to make it concrete.
- Allocate time explicitly. Lena carves out 10 percent of four senior researchers' time, roughly two days a month each, protected on the calendar. Unprotected research time is always the first thing sacrificed to a deadline.
- Set a realistic annual target. Her first-year plan: one peer-reviewed validation study, two white papers for buyers, and one open benchmark dataset. Four artifacts, not forty.
- Build a pipeline with stages. Idea, internal review, experiment, draft, rigor review, publish-or-withhold gate, release. Every project moves through the same stages so nothing skips the checks.
- Put a gate before release. Legal, safety or clinical, and communications sign off using the publish-or-withhold checklist. This is a gate, not a bottleneck; it meets on a fixed cadence so work does not pile up waiting.
- Create incentives. Lena adds "external contribution" to the promotion criteria for senior staff and budgets for two conference trips a year. People do what is rewarded.
- Collaborate outward. One of the four projects is run jointly with a university hospital, which brings data access, credibility, and a co-author who is not on Halden's payroll.
The goal of the first year is not volume. It is to prove the pipeline works end to end, ship a few solid artifacts, and build the muscle, which is why the target is deliberately small enough to hit. A program that sets an ambitious first-year number and misses it tends to be cancelled before anyone learns whether the mechanism was sound. Year two scales what worked and cuts what did not.
Measuring Whether It Is Working
A research program that cannot show its value gets cut in the first budget squeeze, and that squeeze usually arrives before the payoff does. Measure it with a mix of leading indicators, which move early, and lagging indicators, which prove the payoff. Do not judge a first-year program only by citations, which take years to accrue; doing so guarantees the program looks like a failure exactly when it is most vulnerable.
| Leading indicators | Lagging indicators |
|---|---|
| Artifacts shipped against target | Citations and downloads over time |
| Inbound talent mentioning your published work | Design partnerships won that cite the research |
| Invited talks and reviewer requests | Seats on standards or advisory bodies |
| Internal strategy decisions the research changed | Regulatory or buyer trust attributable to it |
The most under-counted measure is the last leading indicator: research that changes your own strategy. When Halden's edge-case benchmark reveals the model underperforms on scans from a particular older machine common in rural clinics, that finding reshapes the product roadmap before it ever earns a citation. Nobody outside the company sees this benefit, and it is the one that would have justified the function even if not a single external reader ever appeared. Knowledge you create about your own systems is often the highest-return knowledge of all.
Lena's First Year, Worked Through
Consider how the whole chapter comes together in Lena's first twelve months. She picks two forms that fit Halden's regulated, trust-driven market: a peer-reviewed validation study and an open edge-case benchmark, plus two buyer-facing white papers as lighter-weight contributions. She protects 10 percent of four researchers' time and runs every project through the pipeline and the publish-or-withhold gate, including the ones nobody expects to be contentious, because a gate that is applied selectively is not a gate.
The numbers, all hypothetical, tell the story. The benchmark is downloaded by a dozen research groups in its first quarter and cited in a hospital network's own model-evaluation guidance. Two candidates in the next senior hiring round name the benchmark as the reason they applied. The validation study, held to a pre-specified protocol, becomes the centerpiece of a design partnership with the university hospital that co-authored it. And internally, the edge-case work redirects roughly a fifth of the next roadmap toward the failure modes the benchmark exposed. Each of those outcomes lands in a different row of the measurement table, which is the point of measuring across several rows rather than waiting on citations.
When Lena returns to the board a year later, she does not repeat the abstract case for research. She shows the partnership, the hires, and the roadmap change, and she reframes the chair's original question. The company was never choosing between buying research and producing it. Producing it was how Halden earned the trust that its buyers, its recruits, and its regulators were quietly demanding all along. That is what knowledge creation looks like as an act of leadership rather than an academic aspiration.
Anti-Patterns
Research programs fail in predictable ways, and most of the failures are decided at setup rather than during the work.
- Equating knowledge creation with writing a paper. Defaulting to the most prestigious form rather than the one that serves the strategy is why many well-intentioned programs stall before shipping anything.
- Unprotected research time. Time that is not carved out and defended on the calendar is the first thing sacrificed to a delivery deadline, every time.
- An ambitious first-year target. Aiming for volume before the pipeline has been proven end to end means missing the number and losing the function before its mechanism was ever tested.
- Deciding publication in the excitement of a result. Without a checklist applied before release, the calls that most need deliberation are the ones made fastest.
- Judging year one by citations. Citations lag by years, so a program measured only that way looks like a failure precisely when its budget is being reviewed.
Practice Prompts
Apply these to your own organization. The exercise that is hardest to complete is usually the one that identifies why the research function has not started.
- Answer the board chair's question for your own company in a few sentences: why should we produce research rather than only buy it? If the answer is generic, it is not yet a case.
- Take the six forms in the table and rank them for your strategy. Name the two or three you would actually pursue and the audience each is aimed at.
- Run a finding your team produced in the last year through the publish-or-withhold checklist. Note which column each of the five questions falls into and whether the answer surprises you.
- Draft a first-year artifact target small enough that you are confident of hitting it, and write down the pipeline stages each artifact must pass through.
Reflection
Think about the knowledge your organization has produced that never left the building: the evaluation nobody wrote up, the failure mode the team understands well and has never described, the internal benchmark that would be useful to others in your field. Ask what has kept it inside. In most cases the barrier is not a competitive or safety judgment, because neither was ever actually made. The work simply never entered a pipeline, and nobody's time was protected to finish it. That is a different problem from secrecy, and it has a different fix.
Then consider the de-risking argument on your own behalf. Pick the assumption your AI strategy most depends on and imagine writing it up for skeptical outsiders with the rigor habits applied: honest baselines, pre-specified evaluation, disclosed limitations, and the full distribution of outcomes rather than the best run. How much of the assumption would survive? Whatever discomfort you feel imagining that exercise is telling you how much unexamined belief is currently loadbearing in your plans.
Glossary
- Knowledge creation: producing new, externally visible knowledge rather than only consuming what others publish; at the visionary level it is part of the leadership role.
- Open benchmark: a published dataset and evaluation protocol that the whole field can use, bought at medium to high effort and paying back in field influence and inbound talent.
- Standards contribution: sustained participation with regulators and standards bodies, whose payoff is shaping the rules you will later operate under.
- Publish-or-withhold checklist: a five-question test covering competitive value, safety and misuse, privacy, rigor, and legal or contractual constraints, where any withhold answer defaults to holding and escalating.
- Reproducibility: the property that someone outside your team can rerun the analysis and obtain the same result, supported by publishing code, permitted data, and exact configurations.
- Pre-specified evaluation: deciding what counts as success before seeing the results, which prevents the goalposts from moving quietly.
- Honest baseline: a comparison against the strongest reasonable alternative rather than a weak strawman chosen to flatter your method.
- Rigor review: the pipeline stage at which the rigor habits are checked, sitting before the publish-or-withhold gate.
- Leading and lagging indicators: measures that move early, such as artifacts shipped and strategy decisions changed, versus measures that prove payoff later, such as citations and partnerships won.
Related Lessons
This chapter pairs closely with Publishing & Writing, which takes over where the publish-or-withhold gate ends and deals with the craft of producing the artifact itself: structuring an argument, writing for a specific audience, and getting work through review. Read in the other direction, International Strategy & Partnerships supplies the context for the collaboration recommendation here, since a joint project with a university hospital is a partnership decision as much as a research one, with the same questions about data access, credit, and shared obligations. Together the three cover why to create knowledge, how to decide what leaves the building, and how to write it well once you have.
Closing
The board chair's question was a good one, and the answer is not that research is prestigious or that serious companies publish. It is that for an organization whose buyers, recruits, and regulators are all making trust decisions about it, producing rigorous public work is the most efficient way to supply the evidence those decisions require, and no amount of research bought from others substitutes for it. The mechanism is unglamorous: pick the two or three forms that fit your strategy, protect a small amount of senior time, run everything through one pipeline with a rigor review and a publish-or-withhold gate, and measure the leading indicators until the lagging ones arrive. Lena's first year produced four artifacts and a changed roadmap. The changed roadmap came from knowledge she created about her own systems, which is the return that arrives first and is counted least.
Key Takeaways
- Creating knowledge buys four things purchasing cannot: influence over the questions the field asks, inbound talent, standing with regulators and standards bodies, and the de-risking that comes from exposing weak assumptions to outside scrutiny.
- The form follows the goal. Papers, white papers, open benchmarks, open-source tools, standards work, and practitioner writing serve different audiences at different costs; pick the two or three that fit your strategy rather than the most prestigious.
- Decide publication with a checklist, not in the excitement of a result. Competitive value, safety and misuse, privacy, rigor, and legal constraints, with holding as the default when any answer lands in the withhold column.
- Rigor is what makes the credibility real. Reproducibility, honest baselines, negative results, pre-specified evaluation, disclosed limitations, and no cherry-picking; in a clinical setting the same habits are a regulatory asset.
- Systems produce research, not intentions. Protected time, a small realistic target, one pipeline with fixed stages, a gate that meets on a cadence, incentives in the promotion criteria, and at least one outward collaboration.
- Measure with leading and lagging indicators together, and never judge a first-year program by citations alone, since they accrue years after the budget decision that would cut the function.
- The highest-return knowledge is often about your own systems. A finding that redirects your roadmap pays back before it earns a single citation.
Frequently Asked Questions
We are not a research organization. Is this relevant to us? The chapter is written for exactly that situation. Halden is a vendor with a product to ship, and the case for creating knowledge rests on commercial outcomes: design partnerships instead of procurement checkboxes, senior candidates who apply because of published work, and regulators who trust organizations that show their work. Nothing here requires an academic identity, only a small protected allocation of senior time and a pipeline.
How do we protect research time when delivery pressure is constant? By making the allocation explicit and putting it on the calendar. Lena's approach is a defined percentage of a small number of senior researchers, protected rather than notional, because unprotected research time is always the first thing sacrificed to a deadline. The second protection is the incentive structure: external contribution in the promotion criteria means the time is not simply generous, it is expected.
What if publishing helps our competitors? That is the first question on the checklist, and it is a real consideration rather than a formality. The test is whether the work builds trust more than it hands rivals an edge, with the withhold condition being that it reveals the core of a durable advantage. Lena holds a paper describing how the flagging threshold is tuned per hospital on exactly those grounds, and publishes an anonymized edge-case benchmark that passes cleanly.
Should we publish results that make our product look bad? Negative and null results are listed as a rigor habit because they are real knowledge and protect others from the same dead end. There is also a self-interested case: the edge-case finding about older machines in rural clinics is unflattering and is also the finding that reshaped the roadmap.
Skill.re