Surya OCR: SOTA model for document intelligence
Datalab has just released Surya OCR, a state-of-the-art OCR model that scores 83.3% on the olmocr benchmark (top under 3B).
650M params
supports 91 languages
5 pages/s on RTX 5090
runs on CPU, GPU, MPS
83.3% olmocr bench score (top under 3B)
Full layout information
Extracts + captions images and diagrams
Strong handwriting, math, form, and table support
Find the GitHub repo here → (don’t forget to star it ⭐️)
Chunked Prefill, clearly explained
Every LLM request has two phases.
Prefill processes the complete input prompt in parallel. It is compute-heavy, builds the request’s KV cache, and produces the first output token.
Decode generates the remaining response one token at a time. Each step repeatedly reads model weights and the KV cache, making it primarily limited by memory bandwidth.
Now imagine Request A is already decoding and streaming tokens when Request B arrives with a 32K-token prompt.
Without chunked prefill, vLLM would process B’s complete prompt in one large pass. Request A cannot generate its next token during that work, so its response appears to freeze.
Chunked prefill splits B’s prompt into smaller token ranges. vLLM can process one chunk, give active requests another opportunity to decode, then continue B’s prefill across later iterations.
The same prompt still gets processed. It simply stops occupying the GPU as one uninterrupted block.
That introduces an important latency tradeoff.
→ Smaller chunks keep active responses moving.
They return control to the scheduler more frequently, reducing inter-token latency spikes for users already receiving an answer.
→ Larger chunks help the new request start sooner.
B cannot generate its first token until every prompt chunk has been processed. Larger chunks complete that work in fewer, more efficient passes, which usually improves time to first token.
→ Chunks that are too small add overhead.
Later chunks must attend to token states created by earlier chunks. Excessive splitting repeatedly reads more of the KV cache and can reduce GPU utilization.
→ Mixed batches can use the GPU more effectively.
Prefill is compute-heavy, while decode is memory-heavy. Scheduling prefill chunks alongside active decodes allows both kinds of work to share an iteration more efficiently.
There is no universal best chunk size.
Long prompts, many concurrent requests, and strict streaming targets usually favor smaller chunks. Workloads that prioritize time to first token may benefit from larger ones.
In vLLM V1, `--max-num-batched-tokens` controls how many tokens can be scheduled in an iteration. It is one of the main settings governing this tradeoff.
Chunked prefill does not reduce the work creatd by a long prompt.
Thanks for reading.



