In today’s newsletter:
Agents without memory aren’t agents at all.
Jev for RAG, clearly explained!
Top 12 agentic use cases for Jev.
Agents without memory aren’t agents at all
An LLM can appear to remember because the application keeps sending previous messages back with each request.
The model itself is still stateless.
Start a new session without stored history, and every preference, decision, and previous outcome disappears.
Agent memory solves this at two different scopes:
1️⃣ Short-term memory
This is the agent’s working state during the current session. It includes recent messages, retrieved documents, uploaded files, tool outputs, and intermediate task state.
2️⃣ Long-term memory
This persists across sessions. It stores information the agent may need again, such as user preferences, known facts, previous outcomes, and task instructions.
Long-term memory can be divided further:
1) Semantic memory stores facts.
For example, a support agent might remember which plan a customer uses or that the customer prefers email over phone calls.
2) Episodic memory stores previous experiences.
This could include how an earlier support issue was resolved or what happened during a previous agent run.
3) Procedural memory stores instructions and learned procedures.
This includes behavioral rules, tool preferences, and steps the agent should follow when completing a task.
These memory types need different write and retrieval policies. A preference should persist until it changes, while an old tool output may only matter during the current session.
The model is not learning through weight updates here. The surrounding system adapts by storing, updating, and retrieving state.
Oracle AI Agent Memory implements this through short-term threads, summaries, durable memories, automatic extraction, scoped retrieval, and context cards.
In Oracle’s documented 80-turn evaluation, the system held input near 1,300 tokens per request while flat history grew past 13,900 tokens by the final turn.
The managed-memory agent also won 48 evaluated turns, while flat history won 13. The remaining 19 were ties.
You can read the full architecture and evaluation here →
Thanks to Oracle for partnering today!
Jev for RAG, clearly explained!
Hybrid search can retrieve a strong shortlist of chunks, but it does not highlight whether each passage contains evidence for the query, only discusses the same topic, or gives the model enough information to answer.
This is important because a passage can rank well because it shares the right terms or sits near the query in embedding space, but it may still be useless as evidence.
Jev adds an explicit evaluation stage between retrieval and generation. BM25 and dense search continue to retrieve candidates. The LLM still writes the answer. Jev scores the candidates before they enter the context window.
This visual explains how you can use Jev to power RAG applications:
Retrieve a broad candidate set
Start with dense and keyword search, then combine their ranked lists with reciprocal rank fusion. The result might be the top 20 passages for a query.
The goal at this stage is recall. You want the candidate set to include the useful evidence, even if it also includes passages that are only loosely related.
Retrieval still sets the upper bound on the system. If the supporting passage is absent from this set, Jev cannot recover it later.
Score every candidate in one request
Pass the user query to Jev as state. For each retrieved passage, define a typed yes-or-no judgment such as:
Does C7 help answer this query
Jev evaluates those judgments together and returns a probability for each candidate. For example:
C1 0.93
C2 0.18
C3 0.76
C4 0.09The packed request is important. A naive implementation makes one model call for every query-passage pair. With 20 candidates, that means 20 calls. Jev can score the full set in one request and return typed outputs that the application can consume directly.
Keep the decision in code
Jev produces probability estimates. The application decides what to do with them.
If the relevance threshold is 0.70, C1 and C3 continue to the generation stage. C2 and C4 never enter the LLM context.
This separation is useful for two reasons. The policy is visible in code, and the team can change it without rewriting the judgment. You can tune the threshold on an evaluation set, inspect false positives and false negatives, and choose the tradeoff that fits the product.
Jev handles the uncertain judgment. Code applies the policy.
Gate the whole answer
Filtering individual passages is only part of the job. A system can retain several relevant passages and still lack enough evidence to answer the query.
The same Jev request can include a second judgment:
Can the query be answered from the retained passages
If that probability falls below the configured threshold, the application can skip generation and return a controlled response such as `not in the documents`.
You can also ask Jev to score whether a candidate contains signs of prompt injection. Treat that score as one filtering input. It is not a security boundary, and it should not replace input isolation, tool permissions, or other controls around the LLM.
The resulting pipeline has clear responsibilities:
Retrieval finds a broad candidate set.
Jev scores which passages contain useful evidence and whether the retained set can support an answer.
Application code applies thresholds and decides whether generation should run.
The LLM writes from the passages that pass those checks.
The useful part of Jev here is the interface. Relevance is no longer an implicit assumption hidden inside a prompt or a ranking score. It becomes a typed probability that the application can log, test, threshold, and audit.
We also built an open-source example that uses Jev as a judge for AI observability with Comet Opik.
It evaluates support traces for groundedness, request coverage, action honesty, and helpfulness, then records the results as an auditable experiment.
Find the Project here on GitHub →
RAG is one of the several use cases that can benefit from Jev. We've covered several more below 👇
Top 12 agentic use cases for Jev:
Jev handles semantic decisions that ordinary code cannot express reliably. It returns typed answers and probabilities, while code continues to cover the workflow.
Here are 12 practical use cases for Jev:
1. Browser next action
Convert the current DOM state into a bounded action such as click, type, or stop. Code executes only valid operation-target pairs. There are already several open-source Jev web agents.
2. Context compaction
Decide which events from a long agent trace should remain. The selected text stays verbatim instead of being replaced with a generated summary.
3. Skill and context loading
Compare the current user turn against the available skills. Load only the instructions needed for that turn instead of filling the context window with every skill.
4. Typed tool-call compilation
Map a natural-language request to a function and fill its typed arguments. Each argument is evaluated separately before code allows execution.
5. Citation verification
Check whether a quoted passage exists and whether the surrounding evidence supports the claim. The output can be supported, unsupported, or contradicted.
6. Extraction verification
Run a cheap extractor first, then use Jev to verify questionable fields. Clean records stay on the fast path while uncertain ones reach a reasoning model.
7. Agent trace evaluation
Turn raw trajectories into queryable labels such as progress and repetition. This avoids asking another LLM to write a full review of every run.
8. Semantic regression tests
Replay a trace suite against a new agent build. Semantic checks can then pass or block prompt, model, tool, and policy changes in CI.
9. Jevgrep code search
Search a codebase by what the code does rather than its exact words. Jev scores candidate snippets and returns the most relevant code first.
10. Entity alignment
Compare two candidate records and decide whether to merge, review, or keep them separate. Candidate generation remains deterministic while Jev handles semantic identity.
11. Retrieval reranking
Let embeddings retrieve a broad candidate set, then use Jev to reorder passages by relevance. The generation model receives the most useful evidence first.
12. Memory promotion gate
Capture a completed agent trace, then judge whether its corrections contain a reusable lesson. Trace-backed lessons can be promoted while task-specific noise is discarded.
If you want to see the final pattern in practice, it is already implemented in the Beacon open-source project.
Beacon captures full sessions across Claude Code, Codex, Cursor, OpenCode, and 20+ agent harnesses, and then Jev identifies which workflows and corrections are worth learning from, so that a lesson discovered by one agent can become available to the others.
GitHub repo: https://github.com/Asymptote-Labs/agent-beacon
(don’t forget to star it ⭐)
Good day!












