>_ mohar@portfolio:~/blog/grey-research-architecture$

← all posts

Grey: A Research Architecture for Auditable Scientific Discovery

2026-07-26 ⏵ 12 min read

The Architecture in Brief

Grey is my agentic research framework for moving beyond LLM-generated speculation toward externally grounded scientific hypothesis generation.

I built it on three pillars.

Geometric bisociation finds cross-domain connections by measuring the embedding space between concepts rather than asking an LLM to invent them.

Simulation-as-judge tests warrants in minimal formal models, replacing LLM self-assigned confidence with simulated trajectory comparisons.

Two-tier knowledge retrieval maintains a curated, human-seeded knowledge graph from PDFs with web search as a flagged fallback.

The unifying ambition is auditable research output where every claim traces to a source, a measurement, or a simulation.

What I Think Works

I started this project because existing research agents share a central weakness: they let the LLM generate hypotheses, justify them, and assign its own confidence. Grey shifts toward externally inspectable signals.

Three ideas I'm genuinely happy with.

Making hypotheses explicit objects through structured argumentation forces the system to distinguish claim from grounds from warrant from backing from qualifier from rebuttal rather than collapsing everything into generated paragraphs.

Separating discovery from interpretation treats the LLM as a semantic compiler over externally generated structure rather than an inventor of unsupported claims.

Provenance as a first-class object with page-level source references, fact-level attribution, and argument-level provenance chains moves beyond citation decoration toward reproducibility.

The PDF-to-knowledge-graph ingestion pipeline with Obsidian-compatible OKF format was a pragmatic choice. The two-tier retrieval design correctly flags web-sourced material as unverified.

Where It Falls Apart

Writing this up, I found a structural weakness that runs through both new pillars.

I replaced an unvalidated LLM opinion with an unvalidated composite of measurements and called the composite validated. Auditable numbers are not the same as valid numbers. The chain of unvalidated assumptions is now longer, just distributed across more components.

Pillar One: Geometric Bisociation

My load-bearing assumption is that linear interpolation between two concept embeddings corresponds to a semantically meaningful region. I stated this as a geometric fact rather than an opinion — and rereading it, I'm no longer comfortable with that framing.

Embedding distance gives a property of a representation learned by a particular encoder. It does not directly give a property of underlying concepts. Semantic embedding spaces are not guaranteed to be globally Euclidean semantic manifolds.

The centroid assumption is especially questionable. Concepts near the midpoint of quantum entanglement and synaptic plasticity in embedding space are not necessarily meaningfully related to both.

Jitter adds further fragility. Isotropic Gaussian noise with an arbitrary scale applied in Euclidean space has no established correspondence to meaningful concept neighborhoods.

Novelty defined as distance from centroid pulls against bridge-worthiness defined as proximity to the bridge region. Nothing arbitrates genuinely surprising from merely irrelevant.

Pillar Two: Simulation-as-Judge

The formalizer extracts mechanism type and domain parameters from warrant text. This reintroduces the exact LLM invention problem I claim to have solved.

Domain parameters like energy functions, interaction matrices, and feedback strengths are not sitting in PDFs. They have to be generated. When the LLM generates them and scipy runs them, the output looks measured but the inputs were guessed.

Two runs of the same equation class with reasonable parameters will tend to produce similar trajectories regardless of whether the underlying warrant is sound. The same functional form applied to both domains creates circularity.

I never specified a null model. Trajectory similarity scores are meaningless without a baseline distribution from unrelated concept pairs run through the same pipeline.

The comparator uses four metrics collapsed into a weighted composite. I asserted the weights. The weighted_composite function is elided. Dynamic time warping is named as correlation but returns a distance, so the sign may silently invert part of the composite.

Unseeded random noise in the simulation templates means reruns will conflate mechanism similarity with noise realization. Reproducibility is aspirational without paired seeding between A and B runs.

Pillar Three: Two-Tier Retrieval

The most defensible pillar still has specific weaknesses.

LLM extraction of concepts from PDFs follows an instruction to extract only what the text says, but that's a prompt, not a guarantee. Extracted paraphrases may not be entailed by the cited segment. I have no entailment check on the extraction step.

Entity resolution by name and alias glosses over cross-domain homonym collisions.

The coverage threshold and flat confidence of 0.3 for Tier 2 facts are unmotivated constants I never justified.

An Internal Inconsistency I Missed

Tier 3 establishes that web search results are untrusted and flagged. But Pillar 1 runs PMI directly against web search and this feeds into the qualifier formula with no tier discount. The ToulminArgument schema has a retrieval_tier field, but the qualifier computation never reads it.

My own trust hierarchy does not apply to half the evidence feeding the measured qualifier.

The Epistemic Problems

Simulation results are presented as qualifiers. But simulation can only tell you what a particular model with particular assumptions and parameters produces. It cannot automatically tell you that the original real-world warrant is true. Those are radically different claims.

The qualifier remains a single number. I moved from "LLM confidence equals 0.82" to "geometry plus evidence plus novelty plus simulation equals 0.82". But the second number is not automatically more meaningful. Arbitrary weights have transformed subjective confidence into unvalidated scoring.

Tier 1 is trusted provenance, not truth. A curated graph can contain incorrect papers, outdated results, flawed experiments, correlation mistaken for causation, or retracted claims. Provenance level is not epistemic truth level.

The formalizer is not just prompt engineering. Translating prose into formal models introduces a huge model-selection bottleneck. I validate the LLM's interpretation of the warrant, not the warrant itself, while believing I validated the warrant.

The ten-model library creates model-selection bias. If a hypothesis requires mechanism seventeen but Grey only knows mechanisms one through ten, the system will find whichever existing model looks most similar.

The Missing Pieces

No experiment establishes whether the geometric system outperforms random bridge discovery.

No comparison shows whether Grey's geometry beats pure LLM hypothesis generation.

No ablation demonstrates whether the knowledge graph improves over retrieval-only systems.

No test confirms whether PMI actually adds useful signal over geometry alone.

No evaluation determines whether simulation improves prediction of expert judgment.

No human baseline exists for how Grey compares with researchers given the same literature.

The evaluation framework measures pipeline structure, not research quality.

What I'd Change

I can reframe the whole architecture around a single thesis. Grey becomes a provenance-preserving hypothesis evolution system that generates cross-domain hypotheses from structured evidence, converts them into testable mechanisms, and updates their epistemic status from computational and empirical evidence.

The research question becomes whether externally grounded hypothesis evolution outperforms LLM-only research agents in generating novel, testable, and subsequently validated scientific hypotheses.

I'd evolve toward hypothesis evolution as central state rather than agent pipeline. Literature retrieval feeds an evidence graph with provenance. Hypothesis generation uses graph bridges, embedding operators, and LLM generation. A hypothesis registry tracks candidates. Mechanism extraction and model selection feed simulation and robustness analysis. Evidence updates and expert review produce auditable reports.

Toulmin becomes one representation for argument rather than the ontology of the entire system. A ResearchObject ontology can encompass claims, evidence, observations, hypotheses, mechanisms, models, simulations, experiments, sources, and arguments.

The qualifier becomes separate dimensions reported rather than collapsed. Evidence high, novelty high, mechanistic support moderate, reproducibility high, external validation unavailable is more scientifically honest than a single number.

Simulation becomes model-conditional mechanistic support rather than a qualifier. Grounding levels become explicit and visible in output. PMI becomes corpus-level association statistics providing one source of statistical support.

How I'd Validate the Architecture

The most important experiment would construct a temporal benchmark. Papers where a later publication discovered something important become test cases. With corpus cutoff at 2015, Grey receives only literature up to that point, generates hypotheses, and future literature from 2016 to 2020 becomes hidden ground truth. Novelty, future validation rate, precision at K, expert plausibility, and time to discovery become measurable.

This tests whether Grey can rediscover relationships that became scientifically accepted later. That is stronger than showing it can produce convincing output.

A cheaper falsification test can validate Pillar 1 alone. Take known bisociative discoveries from the history of science, plus plausible-but-wrong pairs. Run the geometric and PMI pipeline and see whether it separates them. The Swanson ABC linking model from bibliometrics is directly relevant prior art and effectively this PMI-bridge idea already studied.

If this simple test fails, Pillar 2 is being built on a foundation that has not earned it.

What Comes Next

The unambiguous wins carry no epistemic risk. Embedding pre-filtering for the O(n²) connector can ship immediately. Tiered model strategy with strong models for critique can ship immediately. The critic's one-directional bug can be fixed. Streaming output can be added. Error boundaries stay as they are.

The ambitious pillars require validation before they become claims.

The geometric signal is a hypothesis-generation heuristic whose scientific value must be empirically established. Its validity cannot be assumed.

Corpus-level association statistics provide one source of statistical support, not proof of grounding.

Simulation provides model-conditional mechanistic support, not a qualifier.

Formalization is a model-selection and semantic-compilation problem requiring explicit verification, not prompt engineering.

Tier 1 has stronger provenance and curation guarantees, not truth.

The evaluation framework needs negative controls. Random embeddings as a baseline. LLM-only hypothesis generation for comparison. Retrieval-only systems. Geometry-only systems. PMI ablation. Simulation ablation. Human researcher baselines.

The path is clear. Ship the safe components now. Falsify Pillar 1 with known-bridge benchmarks. Build the full architecture with null models and validated weights. Run temporal validation experiments to establish scientific credibility.

Closing

Grey is quite a interesting research project. Some of its measurements are currently dressed-up heuristics. The strongest path is not to add more agents or more simulation templates. It is to make Grey experimentally answer one hard question.

When Grey proposes a cross-domain connection, can I demonstrate that its mechanism is more novel, better grounded, more testable, and more likely to survive future evidence than hypotheses produced by existing LLM research agents?

If I can demonstrate that with proper ablations and forward-validation benchmarks, Grey stops being an impressive personal agent project and starts looking like a legitimate AI-for-scientific-discovery research system.