Michael Spencer
VERIFIED TECHNICAL DOSSIERSources checked

Michael Spencer

Creator & Host of AI Explained

2 Verified ArtifactsSource Checked & Attributed

Verified Proof of Work Artifacts

2 items cataloged

Each artifact below represents an authenticated research publication, production code repository, or technical architectural framework directly authored or co-created by Michael Spencer. Every entry undergoes editorial source verification.

#1
EXPLANATION Checked Sep 22, 2026

AI Explained: Rigorous Evaluation of Frontier Reasoning & Coding Models

Conducted independent empirical benchmarking evaluating reasoning saturation on GPQA, ARC-AGI, and SWE-bench, identifying degradation patterns in extended context retrieval.

Model & Execution Context:GPT-4o, Claude 3.5 Sonnet, o1-preview, SWE-bench Verified, ARC-AGI.
Scope & Limitations

Public benchmarks risk subtle test-set contamination in subsequent foundation model training rounds.

#2
EXPLANATION Checked Sep 22, 2026

Autonomous Software Engineering and Frontier Safety Evaluation Metrics

Formulated comparative methodology for auditing agentic code generation harnesses against real-world GitHub issues, isolating scaffolding efficacy from raw model intelligence.

Model & Execution Context:SWE-bench test harness, Docker execution sandboxes, Claude 3.5 Sonnet.
Scope & Limitations

Unit test pass rates on narrow issue fixes do not evaluate entire codebase architectural coherence.