Not one algorithm but a federation of 1,000+ models across nine surfaces. Reconstructed from system cards, engineering blogs, earnings calls, peer-reviewed audits, and leaks — with every major claim linked to its source.
A from-scratch reimplementation of Logic-LM (EMNLP 2023) that translates natural language reasoning into symbolic logic, executes programs with deterministic solvers, and iteratively refines failures. My notes on what works, what doesn't, and how it compares to the original paper.
My hands-on analysis of the Grey research architecture — where I think it genuinely works, the epistemic gaps I found in its measurement pillars, and the validation path I'd run before trusting any of its numbers.
An offline-first evaluation harness I built to score predicted knowledge-graph triplets against a gold reference set — and what scoring 700 prompts revealed about my fine-tuned Qwen3-0.6B: well-formed output, low fidelity.