TL;DR: People change their mind while they talk. Most voice agents act on the first thing they hear. We put a small layer, the Commit Harness, between the voice model (Gemini 3.8 Live) and its tools. It holds each action until the speaker finishes, replaces it after a correction, cancels it after "never mind", and never runs it twice. On 100 real recordings, our agent passed 67 judged and 55 strict. The same voice model without our changes passed 62 and 50. For the extension, we built an in-car EV assistant that recovers when tools are slow or fail, and an offline Gemma fallback.
The Problem
A person says: "Book a flight to Boston, no, sorry, New York." A voice agent hears the pause after "Boston" and books that flight. Then it hears the correction and books a second flight. The person now has two bookings, and one of them is wrong.
The theme asked for an agent that people can interrupt and correct in real time. Full-Duplex-Bench v3 measures this. It plays 100 recordings of real people who ask for actions out loud, with fillers, pauses, hesitations, false starts and self-corrections. It counts every tool call (an action, such as "track this order") that the agent really runs. One early or extra call fails the recording.
| Domain | Recordings | Example of a tool |
|---|---|---|
| Shopping | 29 | Track an order, add to cart |
| Finance | 25 | Check a balance, pay a bill |
| Housing | 26 | Search listings, book a viewing |
| Travel | 20 | Search flights, book a seat |
The benchmark gives two scores. The strict score counts only tool calls that are exactly right. The judged score lets a language model accept calls that are right in meaning, for example "BOB 12" for "BOB12". The official judge is GPT-4o. We had no access to it, so we used Gemini 2.5 Pro with the benchmark's own judge prompts. We used no OpenAI model anywhere.
The Design: Propose, Settle, Commit
Gemini 3.8 Live listens and speaks. When it wants to act, it proposes a tool call. Our harness takes the proposal and decides when the call can run.
The harness does four things:
- Hold: it waits for 0.9 s of quiet, or 1.8 s when the speaker sounds unsure ("um", "so").
- Replace: after a correction, it drops the held call and keeps the new one.
- Cancel: after "never mind", it drops the call and tells the model that nothing ran.
- Once: it never runs the same call twice, but it keeps a second, different request.
Two deciders tell the harness when the speaker has finished. Reflex reads word patterns such as "no, sorry" and costs nothing. Reasoner (TypeSafe Jev, a small classifier) judges the meaning. If Reasoner takes longer than 0.8 s, Reflex decides alone. We also built a third decider, Listener (Smart Turn v3.2), which hears the tone of the voice. It was wrong too often, so it is off in the submitted run.
We also added two rules around the harness. The prompt rules tell the model that the last value wins and that it must use values exactly as the speaker said them. The identifier rule joins a spelled-out code such as "B-O-B-1-2" into "BOB12" before the tool runs.
Results
| Agent | Judged | Strict | Reply delay |
|---|---|---|---|
| Stock agent (Gemini 3.8 Live, no changes) | 62 | 50 | 3.9 s |
| Ours, submitted | 67 | 55 | 5.3 s |
| Slice (judged) | Stock | Ours |
|---|---|---|
| Shopping, of 29 | 22 | 24 |
| Finance, of 25 | 22 | 22 |
| Housing, of 26 | 5 | 7 |
| Travel, of 20 | 13 | 14 |
| Two requests in one turn, of 18 | 11 | 13 |
| Three requests in one turn, of 16 | 5 | 7 |
| Self-corrections, of 17 | 8 | 7 |
We ran two configurations on 30 September and submitted the better one. The logs of both runs are in the repository. A teammate also ran our setup script on his own laptop and got 53 strict. That is two recordings below our 55, which is a normal difference between runs.
What the Logs Show
The harness logs every decision, so we counted how often it changed what ran. In one full run, Gemini proposed 148 calls. The harness ran 146 of them unchanged, replaced 1 and blocked 1 duplicate. That is 2 of 100 recordings. The reason is simple: Gemini usually waits until it thinks the speaker has finished before it proposes a call. So there is rarely a wrong call left to stop.
Our gain over the stock agent therefore comes from the prompt rules and the identifier rule. We switched them on together, so we cannot say which one helped more. The harness has a cost: it waits about 0.9 s before every call. That is the main reason our agent replies in 5.3 s and the stock agent in 3.9 s.
Ten failures were changes of mind 1.4 to 10.7 s after the first call had already run. No waiting rule can fix those. They need undo, which the extension below provides.
The Extension: an In-Car EV Assistant
The benchmark's tools answer at once and never fail. Real tools are slow and sometimes fail. So we built a recovery layer, a second layer that sits in front of the tools and handles failures. We tested it on an in-car assistant for an electric car, with mock tools (tools that give made-up answers).
One drive, five requests
We played a recorded conversation to the assistant through LiveKit, the same way the benchmark plays its recordings. These are the five requests and what the recovery log shows.
-
"Reroute to the mall, no, sorry, the airport."
One reroute ran, to the airport. 1 call
-
"Find a charging station near downtown."
The search failed twice and worked on the third try. The driver heard only the answer. 2 quiet retries
-
"Book that station for 6 pm." Then: "Book it again."
One booking. The repeat did not run. repeat blocked
-
"Actually, make it 7 pm instead."
The 6 pm booking was cancelled first, then 7 pm was booked. undo, then book
-
"Call roadside assistance." Twice, and the line is down.
After two failed requests, a human took over, with a reference number. hand-off
On real speech
The drive above uses a synthetic voice. To test real voices, we took 11 recordings from SLURP, a public dataset of real people's requests to a home assistant (Bastianelli et al., EMNLP 2020). We picked them by a fixed rule, not by ear. We made the lights tool fail on its first try every time.
The one miss was "light colour for study room". The agent has no colour tool, and it said so.
When the speaker pauses mid-sentence
Next, we put a silence of 1.6 to 3.1 s inside each of the same 11 requests, for example "turn off the ... porch light". The result got much worse.
"Turn the lights off."
Turned the lights on.
"Light colour for study room."
Set the study AC to 22.
"Turn my lights down."
Dimmed them, then rang the phone.
This agent has the recovery layer but not the Commit Harness. The pauses broke it. This is why the two layers must work together in one agent. That is our next step.
An offline fallback with Gemma
If the car loses its connection, a local model must take over. Our teammate Aryan tested two Gemma 4 models with Ollama on a laptop with 32 GB of memory and no graphics card. Each request was a written sentence, the text that the model gets after speech-to-text.
Time per request: about 6 s for 26B, about 2 s for e4b.
The small model was faster, but it refused most real requests and carried out two cancelled actions. After "leave them as they are", it still turned the lights off. So we chose the large model. On our own 40 requests the two models tied. Only the real requests and the cancellations told them apart.
Where It Still Fails
- Late changes of mind: when the person corrects themselves after the call ran, only undo can help. Undo exists only in the extension.
- Pauses inside a request: without the harness, the extension agent took a wrong action in 6 of 11 requests.
- Speed: the harness adds about 0.9 s before each call.
- The judge: our judged scores come from Gemini 2.5 Pro, not GPT-4o. A GPT-4o score can differ by a few recordings.
Try It
You need Linux or WSL, a LiveKit Cloud project and a Gemini API key. The script asks for the keys on the first run and saves them. A full run of 100 recordings takes about 2 hours.
git clone https://github.com/Nithyon/Theme5-Interruptible-Agents
cd Theme5-Interruptible-Agents
./reproduce.sh # the benchmark agent, 100 recordings
./reproduce_extension.sh # offline tests of the recovery layer
./reproduce_extension.sh fallback # the Gemma fallback suite, needs Ollama
Every run that we report has its logs in project-log/runs/, down to each recording.
Credits
Team Reign: Pokala Sai Nithin, Aryan Garg, Lohitashwa and V Preetha, SRM Institute of Science and Technology. The benchmark is Full-Duplex-Bench v3 (arXiv 2604.04847). The real-speech tests use SLURP (Bastianelli et al., EMNLP 2020). AI coding assistants, mostly Claude Code, wrote most of the code and documents. The team made the decisions and is responsible for every claim. The full AI disclosure is in the repository.