In today’s newsletter:
Everyone is sleeping on this new JSON extraction model!
How work is organized inside a GPU.
Jev-style scoring, LLM decoding & structured output.
Everyone is sleeping on this new JSON extraction model!
Pulling structured data out of documents usually takes a chain of OCR, an LLM, and extra code to repair broken JSON.
Lift (open-weights) from datalab does all of this in a single step.
You give Lift a PDF or image along with a JSON schema, and get back a JSON object that matches that schema.
The model reads every page of a document in one pass, so values that span multiple pages still come out right. Any document you can describe in a schema works, from invoices to research papers.
Here’s what you get with lift:
90.2% field accuracy on 11,000 fields
9B model, nearly matches Gemini Flash 3.5 at 3x the speed
built not to hallucinate; returns null if missing
Self-hostable with vLLM or Hugging Face
GitHub repo: https://github.com/datalab-to/lift
(don’t forget to star 🌟)
We will cover this in a hands-on demo soon.
How work is organized inside a GPU
A GPU does not treat a large computation as one job. It keeps dividing that job into smaller pieces until thousands of simple operations can run together.
The easiest way to understand this is to follow the work from the program you launch to the arithmetic the chip performs.
1) A kernel creates the full workload.
A kernel is a function meant to run repeatedly across different pieces of data. When you launch one, it creates a grid. The grid represents all the work required for that launch.
2) The grid is divided into thread blocks.
Each block owns one portion of the workload. Its threads stay together on the same streaming multiprocessor, or SM, so they can coordinate and exchange values through fast shared memory.
The GPU schedules blocks independently. This allows blocks from a large grid to spread across many SMs without needing to coordinate with one another.
3) Each block is divided into warps.
A thread is one lane performing the kernel’s instructions on its own data. The hardware collects threads into fixed groups of 32 called warps.
A block containing 256 threads therefore becomes eight warps. These warps are the units the SM actually chooses between during execution.
All 32 threads in a warp receive the same instruction. They perform it together, but on different values. This is where the GPU gets its width.
It also explains why branching can hurt performance. If threads in one warp choose different paths, the GPU must run each path separately while temporarily disabling the threads that took the other one.
4) Warps execute inside an SM.
An SM contains compute units, warp schedulers, registers, and shared memory. It can keep many warps resident at once, even though only some execute during a given clock tick.
When one warp requests data from slower memory and has to wait, the scheduler selects another ready warp. Switching is extremely cheap because every resident warp already has its state stored on the SM.
The GPU does not eliminate memory delays. It hides them by always having another warp ready to run.
This hierarchy also explains why workload size matters. Too few blocks leave SMs unused. Too few resident warps leave the scheduler with nothing to execute during a memory stall.
The complete path is simple.
A kernel creates a grid → The grid contains blocks → Blocks contain warps → Warps contain 32 threads → Blocks are assigned to SMs → Their warps are scheduled onto compute units.
That structure underpins GPU parallelism, latency hiding, and the need for large batches of similar work.
We wrote the full breakdown on how GPUs work, why they are organized this way, and what that design means for LLM performance.
Jev-style scoring, LLM decoding & structured output
Many LLM requests do not need newly written text.
Routing, classification, policy checks, and ranking usually already have a known set of valid answers, and the only question is which answer best matches the input.
Consider a support ticket:
I was charged twice for the same subscription.
The application must route it to either billing, technical support, or account access.
With normal LLM decoding, the model goes through a full reasoning process. A decoding rule selects one token, appends it to the sequence, and runs the model again.
This continues until the model responds with something like:
Based on the customer’s query, this ticket should go to billing.
The application then has to parse that text to recover the decision it needed from the beginning.
Structured output handles this better, but the underlying process is still generation.
A grammar or schema masks illegal next tokens at every step. The model cannot produce arbitrary prose, but it still generates/decodes the full JSON token by token:
{
“team”: “billing”
}This is useful when the application needs unknown values, nested fields, or tool arguments. The schema guarantees the shape of the response.
Jev-style scoring handles fixed-choice decisions differently.
The application provides the query, the input state, and the allowed answers before inference.
This time, instead of decoding a sentence or JSON object, the model scores the predefined candidates:
billing → 0.90
technical support → 0.10
account access → 0.00The application receives the selected label and its probability distribution directly.
The applicability entirely depends on the downstream use case. They expose different output contracts and perform different types of inference work.
LLM decoding is ideal when the application needs newly written text. It can also return a label, but it must generate and parse that label like any other response.
Structured output decoding is ideal when the values are not known beforehand, but the application requires a predictable schema. The model still generates the response token by token while the decoder prevents invalid structures.
Jev-style scoring is ideal when the valid answers are already known. It cannot write an explanation or produce an unknown field value. Instead, it compares the supplied candidates and returns their probabilities without decoding a sentence or JSON object.
If you want to dive deeper into Jev, we also wrote a full deep dive that explains how system one models like Jev work.
Good day!








