Horace He
VERIFIED TECHNICAL DOSSIERSources checked

Horace He

AI Compiler Architect

2 Verified ArtifactsSource Checked & Attributed

Verified Proof of Work Artifacts

2 items cataloged

Each artifact below represents an authenticated research publication, production code repository, or technical architectural framework directly authored or co-created by Horace He. Every entry undergoes editorial source verification.

#1
RESEARCH Checked Sep 22, 2026

PyTorch 2: Dynamic Python Bytecode Compilation and TorchInductor

Co-authored the ACM ASPLOS 2024 paper detailing TorchDynamo and TorchInductor, which intercept Python frame evaluation bytecodes to capture computational graphs safely and generate fused Triton GPU kernels with 0% code modifications.

Model & Execution Context:TorchDynamo, TorchInductor, OpenAI Triton backend, 160+ benchmark models.
Scope & Limitations

Dynamic control flow with non-tensor Python data structures forces graph breaks that require partial eager execution fallback.

#2
EXPLANATION Checked Sep 22, 2026

Making Deep Learning Go Brrrr: Performance Engineering and GPU Kernel Optimization

Authored the definitive technical guide on GPU architecture, memory hierarchy, arithmetic intensity, and operator fusion, read by hundreds of thousands of AI infrastructure practitioners worldwide.

Model & Execution Context:NVIDIA Ampere/Hopper architectures, SRAM bandwidth, Tensor Cores, roofline analysis.
Scope & Limitations

High-level theoretical roofline models do not account for tail-latency jitter from asynchronous PCIe bus transfers.