EN

Big Tech Bets Voice Will Drive AI Agents

ontime team

1- OpenAI and Google upgraded voice systems to make AI conversations faster, more natural and more context-aware.
2- New models process speech directly, avoiding the text conversion chain used by earlier voice assistants.
3- Rising usage and investment could make speech a primary interface for agents handling extended tasks.

The latest

OpenAI and Google are intensifying efforts to make spoken conversations with chatbots feel natural, betting that voice will overtake typing as AI agents progress from answering questions to carrying out longer-running tasks. Direct speech models promise lower delays and better handling of tone and context, but background noise, reliability, accents and the social discomfort of talking to machines in public remain barriers.

Details

  • Usage signals: Google said live voice usage doubled in the year to April, while voice conversations averaged five times the length of text exchanges. OpenAI said more than 150mn people use ChatGPT voice each week. Microsoft executive Sarah Bird said voice feels natural for assigning long-running tasks and cited colleagues issuing coding commands by phone while working on GitHub.
  • Technical break: Earlier assistants converted speech to text, generated a written answer and then read it aloud. Newer systems process and produce speech directly, reducing the lag, flat delivery and lost context that made previous tools mechanical. OpenAI said its earlier technology had a “lag in intelligence” and faced criticism over slowness and confusion between languages or accents.
  • Product updates: OpenAI released ChatGPT Live last week, describing it as its first voice model in two years. Google’s Leland Rechis said the company adopted a more spontaneous conversational style after finding users became more comfortable once large language models were “behind the microphone”. OpenAI product lead Atty Eleti said speech suits humanlike exchanges, while typing demands more effort.
  • Coding use case: Advocates say speech lets developers explain intent, constraints, uncertainty and desired outcomes more fully than compressed typed prompts. OpenAI’s Sandro Gianella wrote that voice had become important to his use of Codex because speaking encouraged him to explain why he wanted a task completed, rather than immediately issuing an instruction.
  • Hardware response: Silicon Valley start-ups are buying microphones and soundproofing offices to prevent background audio confusing models. Pocket and Plaud offer handheld dictation devices, while Sandbar has designed a ring that takes commands or records notes and responds only to its user’s voice. OpenAI is expected to introduce an ambient assistant with a camera, microphone and speaker early next year, according to people familiar with the project.
  • Funding and flaws: Venture investment in AI voice start-ups reached $7bn in the first quarter, up from just under $1bn a year earlier, according to PitchBook. Tolan technology chief Evan Goldschmidt said background noise, reliability, emotional tone and context still require improvement, describing them as “paper cuts” even as development accelerates.

What’s next

OpenAI’s expected first device early next year is the next concrete test. Usage growth, conversation length and performance across noisy settings, languages and accents will indicate whether voice can move beyond specialist use.

Source

 

What to read next