Samsung PRISM · Theme 05 · Team Reign

A Voice Agent That Waits for What You Finally Meant: 67 of 100 on Full‑Duplex‑Bench v3

TL;DR: People change their mind while they talk. Most voice agents act on the first thing they hear. We put a small layer, the Commit Harness, between the voice model (Gemini 3.8 Live) and its tools. It holds each action until the speaker finishes, replaces it after a correction, cancels it after "never mind", and never runs it twice. On 100 real recordings, our agent passed 67 judged and 55 strict. The same voice model without our changes passed 62 and 50. For the extension, we built an in-car EV assistant that recovers when tools are slow or fail, and an offline Gemma fallback.

The Problem

A person says: "Book a flight to Boston, no, sorry, New York." A voice agent hears the pause after "Boston" and books that flight. Then it hears the correction and books a second flight. The person now has two bookings, and one of them is wrong.

The theme asked for an agent that people can interrupt and correct in real time. Full-Duplex-Bench v3 measures this. It plays 100 recordings of real people who ask for actions out loud, with fillers, pauses, hesitations, false starts and self-corrections. It counts every tool call (an action, such as "track this order") that the agent really runs. One early or extra call fails the recording.

DomainRecordingsExample of a tool
Shopping29Track an order, add to cart
Finance25Check a balance, pay a bill
Housing26Search listings, book a viewing
Travel20Search flights, book a seat

The benchmark gives two scores. The strict score counts only tool calls that are exactly right. The judged score lets a language model accept calls that are right in meaning, for example "BOB 12" for "BOB12". The official judge is GPT-4o. We had no access to it, so we used Gemini 2.5 Pro with the benchmark's own judge prompts. We used no OpenAI model anywhere.

The Design: Propose, Settle, Commit

Gemini 3.8 Live listens and speaks. When it wants to act, it proposes a tool call. Our harness takes the proposal and decides when the call can run.

The path of one request: the voice model proposes a call, the Commit Harness holds it until the turn settles, then the tool runs once. Speaker "Boston, no, NYC" Gemini 3.8 Live listens, proposes Commit Harness Propose Settle Commit hold · replace · cancel · once every decision is logged 12 tools runs once
One request from speech to tool. The harness is the only part we added to the voice path.

The harness does four things:

  1. Hold: it waits for 0.9 s of quiet, or 1.8 s when the speaker sounds unsure ("um", "so").
  2. Replace: after a correction, it drops the held call and keeps the new one.
  3. Cancel: after "never mind", it drops the call and tells the model that nothing ran.
  4. Once: it never runs the same call twice, but it keeps a second, different request.

Two deciders tell the harness when the speaker has finished. Reflex reads word patterns such as "no, sorry" and costs nothing. Reasoner (TypeSafe Jev, a small classifier) judges the meaning. If Reasoner takes longer than 0.8 s, Reflex decides alone. We also built a third decider, Listener (Smart Turn v3.2), which hears the tone of the voice. It was wrong too often, so it is off in the submitted run.

We also added two rules around the harness. The prompt rules tell the model that the last value wins and that it must use values exactly as the speaker said them. The identifier rule joins a spelled-out code such as "B-O-B-1-2" into "BOB12" before the tool runs.

Results

Strict and judged pass counts out of 100 for four runs 020406080 50 62 46 61 55 67 50 64 Stock agent29 Sep run SubmittedSmart Turn on Strict Judged
Passed recordings out of 100. One run each. The judge is Gemini 2.5 Pro with the benchmark's prompts, the same for every agent.
AgentJudgedStrictReply delay
Stock agent (Gemini 3.8 Live, no changes)62503.9 s
Ours, submitted67555.3 s
Slice (judged)StockOurs
Shopping, of 292224
Finance, of 252222
Housing, of 2657
Travel, of 201314
Two requests in one turn, of 181113
Three requests in one turn, of 1657
Self-corrections, of 1787

We ran two configurations on 30 September and submitted the better one. The logs of both runs are in the repository. A teammate also ran our setup script on his own laptop and got 53 strict. That is two recordings below our 55, which is a normal difference between runs.

What the Logs Show

The harness logs every decision, so we counted how often it changed what ran. In one full run, Gemini proposed 148 calls. The harness ran 146 of them unchanged, replaced 1 and blocked 1 duplicate. That is 2 of 100 recordings. The reason is simple: Gemini usually waits until it thinks the speaker has finished before it proposes a call. So there is rarely a wrong call left to stop.

Our gain over the stock agent therefore comes from the prompt rules and the identifier rule. We switched them on together, so we cannot say which one helped more. The harness has a cost: it waits about 0.9 s before every call. That is the main reason our agent replies in 5.3 s and the stock agent in 3.9 s.

Ten failures were changes of mind 1.4 to 10.7 s after the first call had already run. No waiting rule can fix those. They need undo, which the extension below provides.

The Extension: an In-Car EV Assistant

The benchmark's tools answer at once and never fail. Real tools are slow and sometimes fail. So we built a recovery layer, a second layer that sits in front of the tools and handles failures. We tested it on an in-car assistant for an electric car, with mock tools (tools that give made-up answers).

The recovery layer sits between the voice model and the car tools and has five jobs. Gemini 3.8 Live asks for an action Recovery layer Time limit stops an attempt that hangs Retry tries a failed call again, quietly No repeat never books the same thing twice Undo cancels the old booking, then books Hand-off passes to a human after 2 failures Car tools slow, can fail
The five jobs of the recovery layer. The voice model never talks to a tool directly.

One drive, five requests

We played a recorded conversation to the assistant through LiveKit, the same way the benchmark plays its recordings. These are the five requests and what the recovery log shows.

  1. "Reroute to the mall, no, sorry, the airport."

    One reroute ran, to the airport. 1 call

  2. "Find a charging station near downtown."

    The search failed twice and worked on the third try. The driver heard only the answer. 2 quiet retries

  3. "Book that station for 6 pm." Then: "Book it again."

    One booking. The repeat did not run. repeat blocked

  4. "Actually, make it 7 pm instead."

    The 6 pm booking was cancelled first, then 7 pm was booked. undo, then book

  5. "Call roadside assistance." Twice, and the line is down.

    After two failed requests, a human took over, with a reference number. hand-off

On real speech

The drive above uses a synthetic voice. To test real voices, we took 11 recordings from SLURP, a public dataset of real people's requests to a home assistant (Bastianelli et al., EMNLP 2020). We picked them by a fixed rule, not by ear. We made the lights tool fail on its first try every time.

10 / 11requests carried out
8 / 8failures recovered by a retry
1bug found and fixed (the repeat rule)

The one miss was "light colour for study room". The agent has no colour tool, and it said so.

When the speaker pauses mid-sentence

Next, we put a silence of 1.6 to 3.1 s inside each of the same 11 requests, for example "turn off the ... porch light". The result got much worse.

The home assistant on 11 real requests, without and with a pause inside each request 0510 10 0 5 6 Without pausesWith pauses Fully right Wrong or extra action
Requests out of 11, same recordings, one run each.

"Turn the lights off."

Turned the lights on.

"Light colour for study room."

Set the study AC to 22.

"Turn my lights down."

Dimmed them, then rang the phone.

This agent has the recovery layer but not the Commit Harness. The pauses broke it. This is why the two layers must work together in one agent. That is our next step.

An offline fallback with Gemma

If the car loses its connection, a local model must take over. Our teammate Aryan tested two Gemma 4 models with Ollama on a laptop with 32 GB of memory and no graphics card. Each request was a written sentence, the text that the model gets after speech-to-text.

Gemma 4 26B, 15.9 GBGemma 4 e4b, 3.1 GB
Our 40 requests, right
34/36
34/36
Real SLURP light requests, right
45/51
18/51
No fitting tool, did nothing
60/60
60/60
Self-corrections, right
28/29
29/29
"Never mind", did nothing
8/8
6/8

Time per request: about 6 s for 26B, about 2 s for e4b.

The small model was faster, but it refused most real requests and carried out two cancelled actions. After "leave them as they are", it still turned the lights off. So we chose the large model. On our own 40 requests the two models tied. Only the real requests and the cancellations told them apart.

Where It Still Fails

  1. Late changes of mind: when the person corrects themselves after the call ran, only undo can help. Undo exists only in the extension.
  2. Pauses inside a request: without the harness, the extension agent took a wrong action in 6 of 11 requests.
  3. Speed: the harness adds about 0.9 s before each call.
  4. The judge: our judged scores come from Gemini 2.5 Pro, not GPT-4o. A GPT-4o score can differ by a few recordings.

Try It

You need Linux or WSL, a LiveKit Cloud project and a Gemini API key. The script asks for the keys on the first run and saves them. A full run of 100 recordings takes about 2 hours.

git clone https://github.com/Nithyon/Theme5-Interruptible-Agents
cd Theme5-Interruptible-Agents
./reproduce.sh                          # the benchmark agent, 100 recordings
./reproduce_extension.sh                # offline tests of the recovery layer
./reproduce_extension.sh fallback       # the Gemma fallback suite, needs Ollama

Every run that we report has its logs in project-log/runs/, down to each recording.

Credits

Team Reign: Pokala Sai Nithin, Aryan Garg, Lohitashwa and V Preetha, SRM Institute of Science and Technology. The benchmark is Full-Duplex-Bench v3 (arXiv 2604.04847). The real-speech tests use SLURP (Bastianelli et al., EMNLP 2020). AI coding assistants, mostly Claude Code, wrote most of the code and documents. The team made the decisions and is responsible for every claim. The full AI disclosure is in the repository.