What to record when a model ranks memory
An agent memory audit trail for ranked recall: which ranker ordered the results, the egress decision, a hash of the request, and no query or memory text.
By Heartwood MemoryPublished . Dates are Pacific time.
When a model ranks your agent's memory, keep a record for each recall that answers five questions. Which ranker ordered the results? Was anything cleared to leave, and under which policy? Which model was asked, and what request ID came back? What was the request, as a hash? What score did each memory get? If the recall stayed local, record why. That is the core of an agent memory audit trail for ranked recall. Heartwood Memory writes this record for every judged recall, and the judge's record holds no query or memory text.
As of October 1, 2026, Heartwood Memory's Jev judge for recall ranking is available to hosted Team and Professional organizations that opt in; it is off by default.
The record is described in words here on purpose. Its fields are names inside a running system, not a published API, so this post has no schema and nothing to build against yet. What goes into each request is covered on the Jev integration page.
What should an agent memory audit trail record for a ranked recall?
Enough to answer, months later, "why did my agent retrieve that memory, and did anything leave to make it happen?" For each judged recall, Heartwood's judge record holds:
- The ranker. Whether the caller got the model's order or the local ranker's, and whether the model's scores were used at all.
- The setting. Whether the recall ran as a trial or with the judge on.
- The egress decision. The decision itself, its ID, and the provider policy it came from.
- The model and the request ID. Which model version was asked, and the request ID once an answer came back.
- A hash of the request. Not the request.
- The scores. A probability for each memory the model scored and the local ranker's score for each candidate, listed by memory ID.
- The reason for any fallback. Held back, timed out, rate limited, or failed.
- What kind of sensitive content was found. The kind and how many, never the value.
- Timings and token counts for the call.
Why record a hash and not the text?
Because an audit log that copies every query and every memory is a second store of your most sensitive data. It would need its own access rules, and it would have to be cleaned every time a memory is erased.
A hash lets you show that two records refer to the same request, and check a copy you kept against the log, without the log holding the words.
It has limits. In Heartwood the hash is taken over the request as it was prepared, after scrubbing and trimming. You can't rebuild it from the stored memory alone, and you can't read the request back out of it.
Does the record prove what was sent?
Not by the hash alone. Read three things together.
- An egress decision of "allowed" and a hash mean the request was cleared and prepared.
- A request ID and scores mean Jev answered.
- A hash with a fallback reason and no request ID means the request was prepared and no usable answer came back. The reason will be a timeout, a rate limit, an error, or a hold placed by an operator.
In the third case the record can't tell you whether the request reached TypeSafe AI before the deadline. For a privacy review, count it as possibly sent.
Ask this of any tool: is the "what was sent" field written before the send or after it, and can you tell the two apart?
What should the record say when nothing was sent?
It should say so, with the reason. A record of refusals is half of the audit trail. It is how you show a restricted record stayed where it was.
Heartwood records a held-back recall with one of a short list of reasons: a candidate was over the egress ceiling or labeled as personal data, something shaped like a secret was found, the text still looked sensitive after scrubbing, or the egress check refused. Where a pattern matched, the record holds the kind of thing and the number of matches. These records carry no request hash.
Recalls the judge doesn't apply to, such as those from an organization that hasn't opted in, get no judge row. They still get the ordinary recall row described below.
Where does the record live, and how long does it last?
In two places with very different lifetimes.
explain_recall is the short-term view. It explains a recent recall to the session that made it, and on hosted Heartwood a session can explain only its own recalls. It is kept in memory, holds a limited number of recent recalls, and is cleared when the organization's memory changes or the service restarts. Use it to debug today's recall.
The audit log is the copy that lasts. The judge's record is added to Heartwood's hash-chained audit log as its own row. Each row's hash covers the row before it, so an edited row or a row removed from the middle breaks the chain, and the chain can be checked. A chain can't show by itself that the newest rows were cut off the end. That takes an anchor kept somewhere else.
The recall has its own row in the same log, filed under the recall ID, with who asked and how many results were returned. When an egress decision was made, the judge's row is filed under that decision's ID, and each result returned to the caller carries the same ID, so the two can be matched later.
Does the record hold my query or my memories?
The judge's record holds no query or memory text. It holds IDs, scores, hashes, names and counts.
Two things next to it are different, and a reviewer should know both.
The explain_recall entry around the judge's record does hold the query text. Its job is to explain that recall to the session that asked, and it is the short-lived copy.
And "no text" is not "anonymous". The scores are listed by memory ID, and the IDs point at memories in your store. Anyone who can read both the log and the store can look the text up. Set access to the log with that in mind.
What can't this record tell you?
- Whether the ranking was right. A probability is the model's judgment, not a verified answer.
- What TypeSafe AI kept. That follows TypeSafe's terms, at TypeSafe AI legal.
- The words of the request. A hash can't be read back.
- Whether an unanswered request arrived.
How do I trace one recall back?
- Start from the response your agent got. It carries the recall ID, and each result in it names the ranker that ordered it.
- If the local ranker ordered it, read the reason. Either the judge doesn't apply to that recall, or the recall was held back or fell back.
- If Jev ordered it, read the probability for that memory next to the others.
- Read the egress decision and the policy it came from.
- For anything older than a recent session, go to the audit row, and check the chain before you rely on it.
What should I ask of any tool that ranks memory with a model?
- Does each recall record which ranker ordered it, and under which setting?
- Is "what was sent" written before or after the send?
- Is a held-back recall recorded, with a reason?
- Does "no text" describe the whole log or one record in it?
- How long does each copy last, and would an edit or a removed row show?
- Can you match the record to the result the caller actually got?
A trial writes the same record as the live setting. How to use it to compare two rankers is in Jev shadow mode still sends your data. The request itself is covered in What Jev sees when it ranks agent memory.
The hash-chained audit log and explain_recall are part of Heartwood's core and run in every deployment, including self-hosted ones that never call Jev. The quickstart runs governed agent memory on your own machine. If your organization is on a hosted plan and wants the judge, ask to turn on the Jev judge. Checking AI-written memories with Jev is planned and not available yet.
Jev is a model from TypeSafe AI, Inc. Heartwood Memory is made by Edukas Solutions LLC and is not affiliated with or endorsed by TypeSafe AI.
Questions
What should an agent memory audit trail record when a model ranks recall?
Which ranker ordered the results, the egress decision and its policy, the model and request ID, a hash of the request, a score for each memory, and the reason whenever the recall stayed local. Heartwood Memory writes this for every judged recall.
Does Heartwood Memory's judge record contain my query or memory text?
No. The judge's record holds IDs, scores, hashes, names and counts, and no query or memory text. The surrounding explain_recall entry for the same recall does hold the query text, and it is short-lived.
Does a hash in the record prove the request was sent?
Not alone. The hash is written when the request is prepared. A request ID and scores show that Jev answered. A hash with a fallback reason and no request ID means no usable answer came back, and the record can't say whether the request arrived.
Is a held-back recall recorded?
Yes. Heartwood Memory records a recall that was held back together with the reason, and with the kind and number of anything that matched a sensitive pattern, never the value.