Charles Frye
AI Infrastructure Engineer
Sources checkedAI Engineer at Modal Labs and Co-creator of the Full Stack LLM Bootcamp. Ph.D. in Neuroscience from UC Berkeley. Formerly deep learning educator and engineer at Weights & Biases. Specialist in serverless GPU infrastructure, distributed batch inference, containerized vLLM deployment, and practical LLM engineering.
Areas of focus
Professional niches
Proof of Work
The Full Stack LLM Bootcamp: End-to-End Enterprise LLM Production Engineering
Designed and delivered the comprehensive Full Stack LLM Bootcamp, educating tens of thousands of software engineers on prompt design, fine-tuning with LoRA/QLoRA, vector retrieval, evaluation metrics, and cost-efficient deployment.
Rapid evolution in model capabilities requires constant curriculum revisions to address frontier tool-calling interfaces.
Context: PyTorch, Hugging Face, vLLM, LangChain, Modal serverless GPU runtimes.
View missionServerless High-Throughput vLLM Inference Container Orchestration on Modal
Engineered containerized serverless LLM deployment architectures on Modal, orchestrating multi-GPU vLLM inference with cold-start mitigation, dynamic batching, and automated autoscaling from 0 to 100 NVIDIA A100/H100 instances.
Cold-start container image pulling and model weight downloading require warm snapshot caching techniques to maintain sub-second response readiness.
Context: Modal Python SDK, CUDA 12, PagedAttention, Tensor Parallelism across NVLink clusters.
View mission