In today’s newsletter:
Google Advent of Agents Season 3 is here!
Building a production agent harness with LangGraph and Jev.
The LLM scheduling loop, clearly explained.
Google Advent of Agents Season 3 is here!
Advent of Agents Season 3 is a free, 31-day Google Cloud series on building secure and governed AI agents for production.
One hands-on tutorial will be released each day throughout October, covering agent identity, guardrails, sandboxing, observability, evaluations, MCP security, cost controls, and fleet management.
Check it out here: https://adventofagents.com.
Each tutorial starts with a question that customers and security teams actually ask, then answers it with runnable code.
Thanks to Google for partnering today!
Building a production agent harness with LangGraph and Jev
Parts 3 and 4 of our hands-on Agent Engineering course are now available.
Jev is a System One model that returns calibrated probabilities for typed decisions instead of generating open-ended text.
Across these two parts, we use it to route model calls, classify requests, assess tool risk, verify evidence, and detect repeated approaches inside an agent loop.
The complete implementation covers:
Retries, fallback, model routing, and hard call limits
PII redaction, injection filtering, and permission enforcement
Human approval for write actions
Jev-based request classification and tool-risk checks
Structured planning with completion criteria
Evidence checks, Jev probability bands, and independent LLM verification
Repair, replanning, stuck detection, and execution budgets
Memory writes restricted to verified conclusions
Part 3 builds the controls around individual model and tool calls. Part 4 places that agent inside an explicit LangGraph loop:
The reference project adds 46 tests across these controls and includes the complete implementation.
Why care?
Several decisions inside an agent need semantic judgment but do not need another model to generate text.
A router needs to choose from a fixed set of models. A guard needs to decide whether a request should continue. A verifier needs to determine whether the available evidence satisfies one completion criterion.
A traditional LLM can make these decisions, but its output still needs prompting, parsing, and validation. A deterministic rule is cheaper and predictable, but it only catches cases encoded in advance.
Jev handles the middle case. It receives the current state and a typed question, then returns a probability within a fixed schema.
Part 3 uses those probabilities for routing, request screening, and tool-risk checks. Part 4 uses them in a three-layer verifier:
A deterministic evidence check handles clear failures without a model call.
Jev accepts or rejects high-confidence cases.
An independent LLM checker handles the uncertain range.
Jev does not replace permissions, deterministic checks, or human approval. The probability becomes one signal inside a larger control system.
This design is useful for agent decisions that require meaning but have a bounded answer space. It also shows how deterministic code, probabilistic classifiers, and general-purpose LLMs can work together without assigning every decision to the main agent.
By the end of Part 4, the support agent can recover from provider failures, block unauthorized writes, pause actions for approval, verify each investigation step, change approach when progress stops, and return a defined status when the task cannot be completed.
Learn to route, guard, and approve agent actions in Part 3 →
Learn to plan, verify, and stop agent loops correctly in Part 4 →
👉 Over to you: Which bounded agent decisions would you move away from a general-purpose LLM?
The LLM scheduling loop, clearly explained:
Whenever a user sends a prompt to an LLM, the resulting inference request does not go straight from the API endpoint to the GPU.
Instead, it first enters the serving engine, where a scheduler decides when and how its tokens will be processed.
Here is the complete lifecycle:
To dive deeper into the full LLMOps lifecycle, we have covered every bit of this in the LLMOps course, starting from fundamentals to production.
1) The request enters the waiting queue
Each request arrives with a prompt, sampling settings, and a maximum output length.
Prompt lengths vary, and the scheduler cannot know the final generation length in advance. It therefore manages requests one iteration at a time.
2) The scheduler builds the next batch
Before every model step, the scheduler checks two main constraints:
How many tokens can be processed in this iteration
How much KV-cache space remains
It then applies a scheduling policy, often prioritizing running decode requests before admitting new work.
3) New requests enter prefill
During prefill, the model processes the prompt tokens and creates the K and V tensors needed by attention.
This stage can process many prompt tokens in parallel. Long prompts may be split into chunks so they do not block active generations for too long.
4) Running requests enter decode
Once prefill finishes, the request moves to decode.
The model now generates the next token using the KV cache created for all previous tokens. Standard autoregressive decoding usually adds one new token per active request during a model step.
5) The GPU executes the selected work
Modern serving engines can place prefill and decode work in the same GPU batch.
For example, request A may be processing its prompt while requests B and C generate their next tokens.
The batch is therefore not a fixed group that runs until every request finishes. Its composition can change after every model step.
6) The engine updates request state
After execution, generated tokens are sampled and appended to their requests.
A finished request releases its KV-cache blocks and leaves the batch. An unfinished request keeps its state and becomes eligible for the next scheduling iteration:
The scheduler can then use the freed token budget and cache space to admit another waiting request.
In the visual below, C finishes after iteration 1, so D enters during iteration 2. By iteration 3, the active batch contains only A and D.
This iteration-level replacement is continuous batching. It keeps the GPU working while requests with different prompt and output lengths progress independently.
To dive deeper into the full LLMOps lifecycle, we have covered every bit of it in the LLMOps course, starting from fundamentals to production:
Read Part 2 on understanding the core building blocks of LLMs →
Read Part 11 on evaluation of multi-turn systems, tool use evaluations, tracing, and red teaming →
Good day!














