Charles Frye

AI Infrastructure Engineer

Sources checked
ABOUT

AI Engineer at Modal Labs and Co-creator of the Full Stack LLM Bootcamp. Ph.D. in Neuroscience from UC Berkeley. Formerly deep learning educator and engineer at Weights & Biases. Specialist in serverless GPU infrastructure, distributed batch inference, containerized vLLM deployment, and practical LLM engineering.

Areas of focus

Professional niches

THE WORK BEHIND THE PROFILE

Proof of Work

explanationChecked Sep 20, 2026

The Full Stack LLM Bootcamp: End-to-End Enterprise LLM Production Engineering

Designed and delivered the comprehensive Full Stack LLM Bootcamp, educating tens of thousands of software engineers on prompt design, fine-tuning with LoRA/QLoRA, vector retrieval, evaluation metrics, and cost-efficient deployment.

Scope & limitations

Rapid evolution in model capabilities requires constant curriculum revisions to address frontier tool-calling interfaces.

Context: PyTorch, Hugging Face, vLLM, LangChain, Modal serverless GPU runtimes.

View mission
implementationChecked Sep 20, 2026

Serverless High-Throughput vLLM Inference Container Orchestration on Modal

Engineered containerized serverless LLM deployment architectures on Modal, orchestrating multi-GPU vLLM inference with cold-start mitigation, dynamic batching, and automated autoscaling from 0 to 100 NVIDIA A100/H100 instances.

Scope & limitations

Cold-start container image pulling and model weight downloading require warm snapshot caching techniques to maintain sub-second response readiness.

Context: Modal Python SDK, CUDA 12, PagedAttention, Tensor Parallelism across NVLink clusters.

View mission

Guides to evaluating AI expertise