Commit Harness
A voice agent that acts only on what you finally meant.
- Description
- A small layer between the voice model and its tools. It holds each action until you finish speaking, replaces it when you correct yourself, cancels it when you say "never mind", and runs it once.
- Context
- People pause, hesitate and correct themselves while they talk. A voice agent that acts at the first pause books the wrong flight, then books a second one. The benchmark counts every action that really runs, so one early action fails the whole request.
Propose, settle, commit
Gemini 3.8 Live listens and speaks. When it wants to act, it proposes a tool call (an action, such as "track this order"). The harness decides when that call can run. Pick a request below and watch what the harness does.
An illustration of the four rules. The real decisions of every run are in gate_events.log.
Who decides that you finished
Reflex reads word patterns such as "um" or "no, sorry" and costs nothing. It waits for 0.9 s of quiet, or 1.8 s when you sound unsure. Reasoner (TypeSafe Jev, a small classifier) judges the meaning. If Reasoner takes longer than 0.8 s, Reflex decides alone. A third decider, Smart Turn, hears the tone of the voice. It was wrong too often, so it is off in the submitted run.
Two rules around the harness also matter. The prompt rules tell the model that the last value wins. The identifier rule joins "B-O-B-1-2" into "BOB12" before the tool runs.
The benchmark
Full-Duplex-Bench v3 plays 100 recordings of real people who ask for actions out loud, in shopping, finance, housing and travel. The strict score counts only calls that are exactly right. The judged score lets a language model accept calls that are right in meaning. Our judge is Gemini 2.5 Pro with the benchmark's own prompts, in place of GPT-4o, the same for every agent.
Gemini 2.5 Pro, with the benchmark's own judge prompts, unchanged. The official judge is GPT-4o. We had no access to it, and we used no OpenAI model anywhere. Both agents had the same judge. A GPT-4o score can differ by a few recordings.
Passed recordings out of 100. Judge: Gemini 2.5 Pro. One run each. A teammate repeated our run on his own laptop and got 53 strict.
Where the points came from
Judged passes in each slice. The bar shows the share of the slice.
What the logs show
The harness writes every decision to a log, so we counted how often it changed what ran. In one full run, Gemini proposed 148 calls:
Gemini usually waits until it thinks you have finished, so the harness rarely needs to step in. Our gain came from the prompt rules and the identifier rule, which we switched on together. Ten failures were changes of mind after the call had already run. No waiting rule can fix those. They need undo, which the extension below provides.
In the car: when tools fail
The benchmark's tools answer at once and never fail. Real tools do not. So we built a recovery layer for an in-car assistant for an electric car, with mock tools. It tries a failed call again, blocks a repeated booking, cancels an old booking before it makes a new one, and passes you to a human after two failures. Step through the recorded drive:
Played through LiveKit as recorded audio, the way the benchmark plays its recordings. Mock tools, synthetic voice, one run.
On real voices
To test real speech, we used 11 recordings from SLURP, a public dataset of real people's requests to a home assistant (Bastianelli et al., EMNLP 2020). The lights tool failed on its first try every time.
When the speaker pauses mid-sentence
Then we put a silence of 1.6 to 3.1 s inside each of the same requests. This agent has the recovery layer but not the harness, and the pauses broke it.
"Turn the lights off."
2.6 s pause inside
Turned the lights on.
"Light colour for study room."
2.8 s pause inside
Set the study AC to 22.
"Turn my lights down to a lower level of brightness."
3.1 s pause inside
Dimmed them, then rang the phone.
The two layers must work together in one agent. That is our next step.
Offline: when the cloud is gone
If the car loses its connection, a local model must take over. Our teammate Aryan tested two Gemma 4 models with Ollama on a laptop with 32 GB of memory and no graphics card. Each request was a written sentence, the text that the model gets after speech-to-text.
The small model refused most real requests, and after "leave them as they are" it still turned the lights off. We chose the large one.
Try it
You need Linux or WSL, a LiveKit Cloud project and a Gemini API key. The script asks for the keys on the first run. A full run of 100 recordings takes about 2 hours.
git clone https://github.com/Nithyon/Theme5-Interruptible-Agents
cd Theme5-Interruptible-Agents
./reproduce.sh # the benchmark agent
./reproduce_extension.sh # recovery layer tests, no keys
./reproduce_extension.sh fallback # the Gemma fallback, needs Ollama
Every run we report has its logs in project-log/runs/, down to each recording.