
Vipul Ved Prakash
Co-Founder & CEO, Together AI | Former Director of Engineering, Apple | Founder, Topsy & Cloudmark
Verified Proof of Work Artifacts
2 items catalogedEach artifact below represents an authenticated research publication, production code repository, or technical architectural framework directly authored or co-created by Vipul Ved Prakash. Every entry undergoes editorial source verification.
Together Inference Engine & FlashAttention-Optimized Model Serving
Engineered Together AI's custom inference kernel stack, implementing speculative decoding, custom FP8 FlashAttention kernels, and continuous batching to deliver the industry's fastest serving latency for Llama 3, DeepSeek, and Mixtral.
Highly optimized custom kernels require aggressive quantization and continuous hardware profiling for non-standard GPU architectures.
RedPajama: Open Pretraining Datasets for Reproducible Foundation Models
Co-led the open-source release of RedPajama, a 1.2-trillion and 30-trillion token pretraining dataset replicating the LLaMA pretraining corpus to enable transparent, fully reproducible foundation model research.
Web-scale data deduplication and filtering at 30T scale require petabyte-scale distributed compute and risk residual benchmark contamination.