This is a voice agent.

You ask it to book a flight to Boston…

…then you change your mind. "No, sorry, New York."

Most agents act on the first thing they hear.

Ours waits until you finish.

Then it acts. Once.

Samsung PRISM · Theme 05 · Interruptible real-time agents

Commit Harness

A voice agent that acts only on what you finally meant.

Team Reign
  • Pokala Sai NithinTeam lead
  • Aryan GargOffline fallback
  • LohitashwaReproduction, harness rules
  • V PreethaDocumentation & Media Lead
College
SRM Institute of Science and Technology
Voice model
Gemini 3.8 Live
Benchmark
Full-Duplex-Bench v3
100 real recordings
Timeline
29 to 30 Sep 2026
Description
A small layer between the voice model and its tools. It holds each action until you finish speaking, replaces it when you correct yourself, cancels it when you say "never mind", and runs it once.
Context
People pause, hesitate and correct themselves while they talk. A voice agent that acts at the first pause books the wrong flight, then books a second one. The benchmark counts every action that really runs, so one early action fails the whole request.

Propose, settle, commit

Gemini 3.8 Live listens and speaks. When it wants to act, it proposes a tool call (an action, such as "track this order"). The harness decides when that call can run. Pick a request below and watch what the harness does.

    An illustration of the four rules. The real decisions of every run are in gate_events.log.

    Who decides that you finished

    Reflex reads word patterns such as "um" or "no, sorry" and costs nothing. It waits for 0.9 s of quiet, or 1.8 s when you sound unsure. Reasoner (TypeSafe Jev, a small classifier) judges the meaning. If Reasoner takes longer than 0.8 s, Reflex decides alone. A third decider, Smart Turn, hears the tone of the voice. It was wrong too often, so it is off in the submitted run.

    Two rules around the harness also matter. The prompt rules tell the model that the last value wins. The identifier rule joins "B-O-B-1-2" into "BOB12" before the tool runs.

    The benchmark

    Full-Duplex-Bench v3 plays 100 recordings of real people who ask for actions out loud, in shopping, finance, housing and travel. The strict score counts only calls that are exactly right. The judged score lets a language model accept calls that are right in meaning. Our judge is Gemini 2.5 Pro with the benchmark's own prompts, in place of GPT-4o, the same for every agent.

    Our agent, judged
    67 / 100
    55 strict · 5.3 s typical reply
    Gemini 3.8 Live alone, judged
    62 / 100
    50 strict · 3.9 s typical reply
    Judge

    Gemini 2.5 Pro, with the benchmark's own judge prompts, unchanged. The official judge is GPT-4o. We had no access to it, and we used no OpenAI model anywhere. Both agents had the same judge. A GPT-4o score can differ by a few recordings.

    StrictJudged
    Gemini alone
    50
    62
    Our run, 29 Sep
    46
    61
    Submitted
    55
    67
    Smart Turn on
    50
    64

    Passed recordings out of 100. Judge: Gemini 2.5 Pro. One run each. A teammate repeated our run on his own laptop and got 53 strict.

    Where the points came from

    Gemini aloneOur agent
    Shopping (29)
    22
    24
    Finance (25)
    22
    22
    Housing (26)
    5
    7
    Travel (20)
    13
    14
    Three requests in one turn (16)
    5
    7
    Self-corrections (17)
    8
    7

    Judged passes in each slice. The bar shows the share of the slice.

    What the logs show

    The harness writes every decision to a log, so we counted how often it changed what ran. In one full run, Gemini proposed 148 calls:

    146
    ran unchanged
    1
    replaced after a correction
    1
    duplicate blocked
    +0.9 s
    wait before each call

    Gemini usually waits until it thinks you have finished, so the harness rarely needs to step in. Our gain came from the prompt rules and the identifier rule, which we switched on together. Ten failures were changes of mind after the call had already run. No waiting rule can fix those. They need undo, which the extension below provides.

    In the car: when tools fail

    The benchmark's tools answer at once and never fail. Real tools do not. So we built a recovery layer for an in-car assistant for an electric car, with mock tools. It tries a failed call again, blocks a repeated booking, cancels an old booking before it makes a new one, and passes you to a human after two failures. Step through the recorded drive:

    EV assistant · recovery layerrecovery log
    Driver
    Assistant

    Played through LiveKit as recorded audio, the way the benchmark plays its recordings. Mock tools, synthetic voice, one run.

    On real voices

    To test real speech, we used 11 recordings from SLURP, a public dataset of real people's requests to a home assistant (Bastianelli et al., EMNLP 2020). The lights tool failed on its first try every time.

    10/11
    requests carried out
    8/8
    failures recovered by a retry
    1
    bug found and fixed

    When the speaker pauses mid-sentence

    Then we put a silence of 1.6 to 3.1 s inside each of the same requests. This agent has the recovery layer but not the harness, and the pauses broke it.

    10 → 5
    fully right, of 11
    0 → 6
    wrong or extra action, of 11

    "Turn the lights off."

    2.6 s pause inside

    Turned the lights on.

    "Light colour for study room."

    2.8 s pause inside

    Set the study AC to 22.

    "Turn my lights down to a lower level of brightness."

    3.1 s pause inside

    Dimmed them, then rang the phone.

    The two layers must work together in one agent. That is our next step.

    Offline: when the cloud is gone

    If the car loses its connection, a local model must take over. Our teammate Aryan tested two Gemma 4 models with Ollama on a laptop with 32 GB of memory and no graphics card. Each request was a written sentence, the text that the model gets after speech-to-text.

    Gemma 4 26B · 15.9 GB · ~6 sGemma 4 e4b · 3.1 GB · ~2 s
    Our 40 requests
    34/36
    34/36
    Real SLURP light requests
    45/51
    18/51
    No fitting tool, did nothing
    60/60
    60/60
    Self-corrections
    28/29
    29/29
    "Never mind", did nothing
    8/8
    6/8

    The small model refused most real requests, and after "leave them as they are" it still turned the lights off. We chose the large one.

    Try it

    You need Linux or WSL, a LiveKit Cloud project and a Gemini API key. The script asks for the keys on the first run. A full run of 100 recordings takes about 2 hours.

    git clone https://github.com/Nithyon/Theme5-Interruptible-Agents
    cd Theme5-Interruptible-Agents
    ./reproduce.sh                       # the benchmark agent
    ./reproduce_extension.sh             # recovery layer tests, no keys
    ./reproduce_extension.sh fallback    # the Gemma fallback, needs Ollama

    Every run we report has its logs in project-log/runs/, down to each recording.