the open research program

We are testing what the data layer should do before an agent reasons

The evidence now points to a narrow sequence: select for the question, translate the chart’s vocabulary, and traverse only when the task requires a path. Selection held up twice on one real-chart corpus. Fixed microbiology vocabulary passed an untouched holdout. Traversal worked on a constructed path-required task, but not as an automatic addition to the existing star-shaped benchmark. The headline numbers trace to pinned files, and the public manifest states what each result does not prove.

+9.5pp selection vs blunt projection measured in this benchmark 64% of questions overflow 32k raw 61.3% single-LLM judge accuracy — why we use panels
what held up

Build the smallest useful evidence packet.

Question-aware selection beat the frozen blunt projection twice. Fixed microbiology vocabulary improved its registered 44-question holdout stratum. Bounded traversal recovered evidence on a separate task built to require a path.

what did not

More structure is not automatically better.

Pinned references, aggregate summaries, and endpoint reserves were null. QT-4 did not promote traversal. W1 and W2 remain exploratory. A11b exposed a broken normalization choice and did not promote event grouping.

what comes next

Test the compiler, not a graph slogan.

Run stronger retrieval baselines, patient-disjoint terminology and join tests, cross-model and cross-server cells, then the governed Bonfire product benchmark. Graph storage remains an implementation question, not a research result.

Run it yourself — five tools, zero upload

These tools explain or reproduce parts of the research. Some are concept demos, not measured experiment arms. Everything computes in your browser; nothing leaves your machine.

The next claim has to survive the next test.

The research supports a context-compiler direction. It does not yet validate Bonfire, graph storage, or generality across models, servers, and institutions.