LLM capabilities, model comparisons, practical applications, technical deep dives
Who this is for: AI researchers, developers, product builders, tech enthusiasts
20 articles.
DPO vs. RLHF: The Real Story on What Works for LLM Alignment #60— The debate between DPO and RLHF is full of misinformation. Having implemented both at scale, I'm cutting through the noise to give you the unvarnished truth about which alignment technique is right for your model and when.
The Truth About LLM Inference Costs: A Data-Driven Analysis #13— I analyzed the inference costs of the top 20 LLMs and the results will surprise you. This deep dive reveals the hidden factors that drive up costs and provides a framework for making smarter, more economical choices for your AI stack.
The Rise of Small Language Models: Why Bigger Isn't Always Better #38— Everyone is obsessed with massive, billion-parameter models, but they're missing the bigger picture. I'll show you why small language models are the future of AI and how they're quietly powering a revolution in efficiency and accessibility.
On-Device AI: The Ultimate Guide to Running LLMs Locally— The cloud is expensive and slow. I've spent the last two years focused on on-device AI, and I'm sharing everything I've learned about running powerful LLMs directly on your users' hardware. This is the future of private, personalized AI.
The Rise of Small Language Models: Why Bigger Isn't Always Better #24— Everyone is obsessed with massive, billion-parameter models, but they're missing the bigger picture. I'll show you why small language models are the future of AI and how they're quietly powering a revolution in efficiency and accessibility.
11 Things I Learned About Token Economics After Analyzing 100+ LLMs #11— I went down the rabbit hole of token economics, analyzing over 100 different large language models. The results were shocking. Here are the 11 most critical lessons I learned about pricing, efficiency, and the future of tokenization.
On-Device AI: The Ultimate Guide to Running LLMs Locally— The cloud is expensive and slow. I've spent the last two years focused on on-device AI, and I'm sharing everything I've learned about running powerful LLMs directly on your users' hardware. This is the future of private, personalized AI.
Model Merging: The Secret Weapon for Creating Hyper-Specialized LLMs— Why train a model from scratch when you can merge the best of what's already out there? I'll walk you through the art and science of model merging, a powerful technique for creating highly specialized models with a fraction of the effort.
The Truth About LLM Inference Costs: A Data-Driven Analysis #12— I analyzed the inference costs of the top 20 LLMs and the results will surprise you. This deep dive reveals the hidden factors that drive up costs and provides a framework for making smarter, more economical choices for your AI stack.
I Spent 5 Years Optimizing LLM Inference, Here's The Truth Nobody Talks About— After half a decade in the trenches of inference optimization, I'm sharing the counterintuitive lessons I learned that challenge everything you think you know. This isn't about chasing benchmarks; it's about real-world performance and the surprising trade-offs nobody mentions.
Why Most Founders Get RLHF Completely Wrong (And Fix It)— I've seen countless startups burn through cash trying to implement RLHF without understanding the fundamentals. Here's the painful truth about why it fails and a simple framework for getting it right from day one.
How to Build Multimodal LLMs That Don't Suck (A Practical Guide)— Building multimodal LLMs is the next frontier, but most attempts are clunky and impractical. I'm sharing my playbook for creating seamless, intuitive multimodal experiences that users will actually love, based on my experience shipping three of them.
11 Things I Learned About Token Economics After Analyzing 100+ LLMs— I went down the rabbit hole of token economics, analyzing over 100 different large language models. The results were shocking. Here are the 11 most critical lessons I learned about pricing, efficiency, and the future of tokenization.
Why Most Founders Get RLHF Completely Wrong (And How to Fix It)— I've seen countless startups burn through cash trying to implement RLHF without understanding the fundamentals. Here's the painful truth about why it fails and a simple framework for getting it right from day one.
What I Learned After 5 Years of Optimizing LLM Inference— After working in inference optimization for five years, I'm sharing the unexpected lessons I've picked up. This is about practical results and the trade-offs most people don’t talk about.
Model Merging: The Secret Weapon for Creating Hyper-Specialized LLMs— Why train a model from scratch when you can merge the best of what's already out there? I'll walk you through the art and science of model merging, a powerful technique for creating highly specialized models with a fraction of the effort.
The Counterintuitive Guide to Model Distillation That Actually Works #36— I distilled my first model in 2021 and failed miserably. After years of trial and error, I've developed a counterintuitive approach to model distillation that delivers smaller, faster models without sacrificing performance. Here's my step-by-step process.
The Counterintuitive Guide to Model Distillation That Actually Works #20— I distilled my first model in 2021 and failed miserably. After years of trial and error, I've developed a counterintuitive approach to model distillation that delivers smaller, faster models without sacrificing performance. Here's my step-by-step process.
DPO vs. RLHF: The Real Story on What Works for LLM Alignment #25— The debate between DPO and RLHF is full of misinformation. Having implemented both at scale, I'm cutting through the noise to give you the unvarnished truth about which alignment technique is right for your model and when.
How to Build Multimodal LLMs That Don't Suck (A Practical Guide) #16— Building multimodal LLMs is the next frontier, but most attempts are clunky and impractical. I'm sharing my playbook for creating seamless, intuitive multimodal experiences that users will actually love, based on my experience shipping three of them.