Daily Dose of Data Science
Subscribe
Sign in
Home
Sponsor
Premium
Archive
Leaderboard
About
Latest
Top
Discussions
KV vs Prefix vs Prompt vs Semantic Caching
...explained with best practices in production.
11 hrs ago
•
Avi Chawla
7
[Hands-on] Turn Scientific Figures Into Structured Data with Mistral OCR
A full walkthrough of the extraction schema, with code.
Aug 26
•
Avi Chawla
9
Build a Multi-Agent GTM Intelligence System
...explained with code.
Aug 25
•
Avi Chawla
8
1
Preloading Knowledge Into a Model Instead of Retrieving It
How to process your corpus once, skip retrieval entirely, and serve every query from a stored cache. Three parts covering the full spectrum.
Aug 24
•
Avi Chawla
6
How Semantic Code Navigation Cuts Agent Token Costs by up to 36%
Understanding what an agent actually does with the tokens before it writes code.
Aug 21
•
Avi Chawla
11
1
What is (was?) GIL in Python?
...explained with code.
Aug 20
•
Avi Chawla
14
Kimi K3's Sandbox Problem Finally Has an Open-Source Fix
...explained with code.
Aug 19
•
Avi Chawla
5
1
Grok Bot Masterclass
Everything you need to understand, set up, and get real work out of Grok Bot.
Aug 18
•
Avi Chawla
8
How a GPU Actually Works
The intuition an LLM engineer needs. Understand techniques like quantization, speculative decoding, and continuous batching in one place.
Aug 17
•
Avi Chawla
11
A Cheaper Model Does Not Imply a Cheaper Turn
The practical implications of model routing, clearly explained.
Aug 16
•
Avi Chawla
15
1
1
How Production LLMs Reason Better At Inference Time
8 techniques, explained visually.
Aug 14
•
Avi Chawla
9
1
Continuous Batching in LLMs
The technique behind vLLM's 23x throughput jump and the default scheduler in every serving engine.
Aug 13
•
Avi Chawla
15
1
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts