Michael Spencer
Creator & Host of AI Explained
Verified Proof of Work Artifacts
2 items catalogedEach artifact below represents an authenticated research publication, production code repository, or technical architectural framework directly authored or co-created by Michael Spencer. Every entry undergoes editorial source verification.
AI Explained: Rigorous Evaluation of Frontier Reasoning & Coding Models
Conducted independent empirical benchmarking evaluating reasoning saturation on GPQA, ARC-AGI, and SWE-bench, identifying degradation patterns in extended context retrieval.
Public benchmarks risk subtle test-set contamination in subsequent foundation model training rounds.
Autonomous Software Engineering and Frontier Safety Evaluation Metrics
Formulated comparative methodology for auditing agentic code generation harnesses against real-world GitHub issues, isolating scaffolding efficacy from raw model intelligence.
Unit test pass rates on narrow issue fixes do not evaluate entire codebase architectural coherence.