Eugene Yan
VERIFIED TECHNICAL DOSSIERSources checked

Eugene Yan

Applied Scientist, Amazon | Enterprise RAG & Recommendation Architect

2 Verified ArtifactsSource Checked & Attributed

Verified Proof of Work Artifacts

2 items cataloged

Each artifact below represents an authenticated research publication, production code repository, or technical architectural framework directly authored or co-created by Eugene Yan. Every entry undergoes editorial source verification.

#1
IMPLEMENTATION Checked Sep 20, 2026

Amazon: Scalable Model Serving Architecture & Latency Optimization

A production infrastructure blueprint developed by Eugene Yan at Amazon, implementing continuous batching, quantized weights, and horizontal autoscaling for high-concurrency model inference.

Model & Execution Context:Optimized for high-throughput PyTorch / vLLM runtime serving.
Scope & Limitations

Deployment specifications are tailored to modern GPU cluster infrastructure; requires containerized execution runtimes.

#2
EXPLANATION Checked Sep 20, 2026

Patterns for Building LLM-based Systems & Products

Comprehensive reference manual documenting production patterns for evaluation, guardrails, retrieval-augmented generation, and fine-tuning with cost/latency trade-offs.

Model & Execution Context:Python, vector indexes, cross-encoders, and LLM evaluation frameworks.
Scope & Limitations

System design patterns; individual implementations must be profiled for specific latency constraints.