the open research program
We are testing what the data layer should do before an agent reasons
The evidence now points to a narrow sequence: select for the question, translate the chart’s vocabulary, and traverse only when the task requires a path. Selection held up twice on one real-chart corpus. Fixed microbiology vocabulary passed an untouched holdout. Traversal worked on a constructed path-required task, but not as an automatic addition to the existing star-shaped benchmark. The headline numbers trace to pinned files, and the public manifest states what each result does not prove.
Build the smallest useful evidence packet.
Question-aware selection beat the frozen blunt projection twice. Fixed microbiology vocabulary improved its registered 44-question holdout stratum. Bounded traversal recovered evidence on a separate task built to require a path.
More structure is not automatically better.
Pinned references, aggregate summaries, and endpoint reserves were null. QT-4 did not promote traversal. W1 and W2 remain exploratory. A11b exposed a broken normalization choice and did not promote event grouping.
Test the compiler, not a graph slogan.
Run stronger retrieval baselines, patient-disjoint terminology and join tests, cross-model and cross-server cells, then the governed Bonfire product benchmark. Graph storage remains an implementation question, not a research result.
Query-aware clinical context
The headline results: question-only selection beat frozen blunt projection +9.5pp (preregistered, p=1.3×10⁻⁵) at 43% less payload; QT-4 confirmed microbiology vocabulary on a fresh 374-question holdout.
Open → the lab notebookFindings, nulls, and open tests
The current experiment ledger: confirmed findings, exploratory signals, grading limitations, nulls, and the next tests needed.
Open → the full write-upClinical context engineering for FHIR agents
The complete report: methods, failure forensics, statistics, corrections, and what each finding does and does not license.
Open → the explainer · part oneThe Agent Walks the Graph
Watch one clinical question traverse a FHIR reference graph until the context window overflows — every response fits, the conversation doesn’t, and the answer was on page 12.
Open → the explainer · part two…and Lives
Same graph, same window: the data layer does the traversal, one bounded cited packet enters the context, and the page-12 value arrives with a citation. This is an illustrative design, not a measured Bonfire result.
Open →Run it yourself — five tools, zero upload
These tools explain or reproduce parts of the research. Some are concept demos, not measured experiment arms. Everything computes in your browser; nothing leaves your machine.
The next claim has to survive the next test.
The research supports a context-compiler direction. It does not yet validate Bonfire, graph storage, or generality across models, servers, and institutions.