DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents
Original reporting by arXiv (cs.AI)

Full-duplex voice agents are advanced AI systems designed to engage in natural, real-time spoken conversations, continuously managing turns, backchannels, and interruptions with human-like fluidity. While such agents are increasingly deployed in roles requiring nuanced social understanding, existing benchmarks often evaluate them on explicit turn-management instructions. This creates a disconnect with real-world scenarios, where agents are expected to infer appropriate conversational behavior from assigned roles or personas. To bridge this gap, new research introduces DuplexSpeechBench-IFEval (DSB-IFEval), a comprehensive benchmark for evaluating implicit instruction-following in real-time spoken interaction.
DSB-IFEval comprises 1,038 test cases across eight diverse assistant roles, assessing five conditioning protocols, from default behavior to persona-implied actions and even instruction conflicts. It uses a deterministic Instruction Adherence Score (IAS) for real-time floor management and an LLM-judged Persona Adherence Score (PAS) for content consistency.
Performance Divergence
The evaluation across six real-time speech systems revealed significant architecture-dependent trade-offs. Full-duplex models like F-Actor and PersonaPlex proved highly sensitive to whether behavioral instructions were explicit or merely implied by a persona, with adherence in floor management dropping under persona-only conditions. Conversely, systems such as GPT-Realtime and MiniCPM-o demonstrated strong adherence to persona-consistent *content* but struggled to adapt their floor behavior across different instruction types. Critically, agents also exhibited difficulty overriding conflicting directives, particularly when safety issues were involved. These findings underscore that inferring role-based behavior, executing it contextually, and resolving competing instructions remain distinct and formidable challenges for the next generation of full-duplex voice agents.
The introduction of DuplexSpeechBench-IFEval (DSB-IFEval) marks a crucial step forward in evaluating full-duplex voice agents, moving beyond explicit turn-management to assess their ability to infer and act on implicit instructions derived from personas and roles. This comprehensive benchmark reveals that despite significant advancements in conversational AI, current systems face nuanced and architecture-dependent challenges. While some full-duplex models demonstrate strong adherence to explicit directives, their performance in inferring behavior from persona-only conditioning significantly diminishes. Conversely, other leading models excel at persona-consistent content but struggle to adapt their real-time floor management or execute proactive actions based on inferred or even explicit behavioral instructions. The findings underscore that the complexities of inferring appropriate conversational behavior, executing it synchronously, and resolving conflicting instructions remain distinct, unresolved hurdles for AI systems striving for truly natural human-like interaction.
Towards natural interaction These insights carry significant implications for the future of AI. As voice agents become more integrated into daily life, their capacity to understand subtle social cues and adapt their conversational style implicitly is paramount for intuitive and trustworthy interaction. DSB-IFEval highlights that simply generating persona-consistent content is insufficient; models must also exhibit persona-consistent behavior in real-time, including intelligent turn-taking, backchanneling, and interruption management. The benchmark further reveals a critical gap in handling safety conflicts, indicating that even when systems can navigate conflicting directives, overriding them for user safety remains a struggle. This necessitates a renewed focus on developing AI architectures capable of robust, context-aware reasoning for conversational dynamics and ethical decision-making, providing a clear roadmap for engineers to build more sophisticated, reliable, and ultimately, more human-centric conversational experiences.
Frequently asked questions
- What is DuplexSpeechBench-IFEval, and what does it evaluate in AI voice agents?
- DSB-IFEval is a novel benchmark designed to assess full-duplex voice agents' ability to follow implicit conversational instructions, particularly those inferred from assigned roles or personas. It evaluates how agents manage real-time interactions, including when to listen, interrupt, or yield the floor, and ensures content remains consistent with their defined character, moving beyond explicit turn-management commands.
- What are the main challenges for full-duplex AI voice agents in natural conversations?
- Full-duplex voice agents face challenges in inferring appropriate conversational behaviors from implicit roles, executing them contextually, and resolving competing instructions. Models show trade-offs; some adapt floor management poorly when behavior is persona-implied versus explicitly stated, while others maintain persona-consistent content but struggle with proactive actions. Overriding conflicting directives, especially safety-related ones, also remains a significant difficulty.
- How do different full-duplex AI models perform when following implicit conversational instructions?
- Performance varies among full-duplex AI models. Full-duplex models like F-Actor and PersonaPlex are more sensitive to whether conversational behavior is explicitly stated or inferred from a persona, showing reduced adherence under persona-only conditioning for floor management. In contrast, models such as GPT-Realtime prioritize persona-consistent content but often struggle to adapt their floor behavior across different instruction types.