This evaluation probes the accuracy-latency-efficiency trade-off of streaming voice agents under realistic ASR conditions. It measures how well a system maintains reasoning quality while minimizing computational overhead and response delays when processing natural speech with disfluencies, misrecognitions, and non-uniform speaking rates. Use when the user wants to benchmark on VERA (AIME and GPQA-Diamond), Spoken-MQA, BigBenchAudio, Pause-and-Repair Benchmark, or asks about evaluating this ta...
Scanned 9/11/2026
Install to Claude Code
npx -y skills add qhjqhj00/research-skills-pool --skill lts-voiceagent-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Lts Voiceagent Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/qhjqhj00-lts-voiceagent-eval)More formats (shields.io, HTML) on the badges page.
---
name: lts-voiceagent-eval
description: This evaluation probes the accuracy-latency-efficiency trade-off of streaming voice agents under realistic ASR conditions. It measures how well a system maintains reasoning quality while minimizing computational overhead and response delays when processing natural speech with disfluencies, misrecognitions, and non-uniform speaking rates. Use when the user wants to benchmark on VERA (AIME and GPQA-Diamond), Spoken-MQA, BigBenchAudio, Pause-and-Repair Benchmark, or asks about evaluating this task. Reports Accuracy.
metadata:
skill_kind: dataset_eval
source_arxiv: 2601.19952
bibtex_key: zou2026ltsvoiceagent
confidence: high
---
# lts-voiceagent-eval
> LTS-VoiceAgent: A Listen-Think-Speak Framework for Efficient Streaming Voice Interaction via Semantic Triggering and Incremental Reasoning — Wenhao Zou et al. (arXiv:2601.19952, 2026)
## What this evaluates
This evaluation probes the accuracy-latency-efficiency trade-off of streaming voice agents under realistic ASR conditions. It measures how well a system maintains reasoning quality while minimizing computational overhead and response delays when processing natural speech with disfluencies, misrecognitions, and non-uniform speaking rates.
## Datasets
- **VERA (AIME and GPQA-Diamond)** — total ?; splits: test (-1)
- **Spoken-MQA** — total ?; splits: test (-1)
- **BigBenchAudio** — total ?; splits: test (-1)
- **Pause-and-Repair Benchmark** — total ?; splits: test (-1)
## Metrics
- `Accuracy` **(primary)** — range: percent
- Standard exact-match or pass@1 accuracy on question-answering tasks, measuring the percentage of correctly answered queries.
- `Time-to-First-Sentence (TTFS)` — range: other
- Wall-clock time from the start of audio input to the generation of the first complete sentence.
- `Number of Forward-pass Evaluations (NFE)` — range: other
- Total count of LLM forward-pass executions performed during the inference process.
- `Number of Interruption Times (NIT)` — range: other
- Count of mid-reasoning interruptions triggered by subsequent user speech or ASR revisions.
- `Interruption Rate` — range: percent
- NIT divided by NFE, representing the fraction of forward passes that are interrupted by subsequent speech.
## Input / output format
**Input**: Streaming audio processed by a unified internal ASR API at 200ms granularity (16kHz mono PCM, 3,200 samples per chunk), yielding incremental text transcripts fed to the LLM.
**Output**: Streaming text response generated via greedy decoding (temperature=0, max 4,096 new tokens) from the Qwen3-8B backbone.
## Scoring recipe
```python
def compute_metrics(predictions, golds, nfe, nit):
correct = sum(1 for pred, gold in zip(predictions, golds) if pred.strip() == gold.strip())
accuracy = (correct / len(golds)) * 100
interruption_rate = (nit / nfe) * 100
return {'accuracy': accuracy, 'interruption_rate': interruption_rate}
# TTFS and NFE are recorded directly from system logs during inference
```
## Common pitfalls
- Evaluating the semantic trigger classifier in isolation rather than end-to-end, which fails to capture ASR jitter and the non-unique nature of spoken semantic boundaries.
- Comparing against baselines that use fixed text chunks or heuristic noise instead of realistic streaming ASR, leading to inflated efficiency claims.
- Confusing Time to First Token (TTFT) with Time-to-First-Sentence (TTFS), which better reflects perceived responsiveness in interactive voice agents.
## Evidence (verbatim from paper)
> We evaluate quality via Accuracy, latency via Time to First Token and Time-to-First-Sentence (TTFS), and efficiency via the Number of Forward-pass Evaluations (NFE) and the Number of Interruption Times (NIT). To quantify trigger validity, we report the Interruption Rate, defined as $ ext{NIT}/ ext{NFE}$; lower values indicate fewer mid-reasoning interruptions caused by subsequent speech and therefore more precise trigger timing.
## Citation
```bibtex
@misc{zou2026ltsvoiceagent,
title={LTS-VoiceAgent: A Listen-Think-Speak Framework for Efficient Streaming Voice Interaction via Semantic Triggering and Incremental Reasoning},
author={Wenhao Zou et al.},
year={2026},
note={arXiv:2601.19952}
}
```
- arXiv: 2601.19952
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!