Thursday, May 22, 2025
LBNN
  • Business
  • Markets
  • Politics
  • Crypto
  • Finance
  • Energy
  • Technology
  • Taxes
  • Creator Economy
  • Wealth Management
  • Documentaries
No Result
View All Result
LBNN

AI-powered headphones offer group translation with voice cloning and 3D spatial audio

Simon Osuji by Simon Osuji
May 10, 2025
in Artificial Intelligence
0
AI-powered headphones offer group translation with voice cloning and 3D spatial audio
0
SHARES
0
VIEWS
Share on FacebookShare on Twitter


AI headphones translate multiple speakers at once, cloning their voices in 3D sound
Credit: University of Washington

Tuochao Chen, a University of Washington doctoral student, recently toured a museum in Mexico. Chen doesn’t speak Spanish, so he ran a translation app on his phone and pointed the microphone at the tour guide. But even in a museum’s relative quiet, the surrounding noise was too much. The resulting text was useless.

Related posts

13 Best Memorial Day Sales on Our Favorite Gear (2025)

13 Best Memorial Day Sales on Our Favorite Gear (2025)

May 22, 2025
The Enhanced Games Has a Date, a Host City, and a Drug-Fueled World Record

The Enhanced Games Has a Date, a Host City, and a Drug-Fueled World Record

May 21, 2025

Various technologies have emerged lately promising fluent translation, but none of these solved Chen’s problem of public spaces. Meta’s new glasses, for instance, function only with an isolated speaker; they play an automated voice translation after the speaker finishes.

Now, Chen and a team of UW researchers have designed a headphone system that translates several speakers at once, while preserving the direction and qualities of people’s voices. The team built the system, called Spatial Speech Translation, with off-the-shelf noise-canceling headphones fitted with microphones. The team’s algorithms separate out the different speakers in a space and follow them as they move, translate their speech and play it back with a 2-4 second delay.







University of Washington researchers designed a headphone system that translates several people speaking at once, following them as they move and preserving the direction and qualities of their voices. The team built the system, called Spatial Speech Translation, with off-the-shelf noise-cancelling headphones fitted with microphones. Credit: Chen et al./CHI ’25

The team presented its research Apr. 30 at the ACM CHI Conference on Human Factors in Computing Systems in Yokohama, Japan. The code for the proof-of-concept device is available for others to build on. “Other translation tech is built on the assumption that only one person is speaking,” said senior author Shyam Gollakota, a UW professor in the Paul G. Allen School of Computer Science & Engineering. “But in the real world, you can’t have just one robotic voice talking for multiple people in a room. For the first time, we’ve preserved the sound of each person’s voice and the direction it’s coming from.”

The system makes three innovations. First, when turned on, it immediately detects how many speakers are in an indoor or outdoor space.

“Our algorithms work a little like radar,” said lead author Chen, a UW doctoral student in the Allen School. “So they’re scanning the space in 360 degrees and constantly determining and updating whether there’s one person or six or seven.”

The system then translates the speech and maintains the expressive qualities and volume of each speaker’s voice while running on a device, such mobile devices with an Apple M2 chip like laptops and Apple Vision Pro. (The team avoided using cloud computing because of the privacy concerns with voice cloning.) Finally, when speakers move their heads, the system continues to track the direction and qualities of their voices as they change.

The system functioned when tested in 10 indoor and outdoor settings. And in a 29-participant test, the users preferred the system over models that didn’t track speakers through space.

In a separate user test, most participants preferred a delay of 3-4 seconds, since the system made more errors when translating with a delay of 1-2 seconds. The team is working to reduce the speed of translation in future iterations. The system currently only works on commonplace speech, not specialized language such as technical jargon. For this paper, the team worked with Spanish, German and French—but previous work on translation models has shown they can be trained to translate around 100 languages.

“This is a step toward breaking down the language barriers between cultures,” Chen said. “So if I’m walking down the street in Mexico, even though I don’t speak Spanish, I can translate all the people’s voices and know who said what.”

Qirui Wang, a research intern at HydroX AI and a UW undergraduate in the Allen School while completing this research, and Runlin He, a UW doctoral student in the Allen School, are also co-authors on this paper.

More information:
Tuochao Chen et al, Spatial Speech Translation: Translating Across Space With Binaural Hearables, Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (2025). DOI: 10.1145/3706598.3713745

Provided by
University of Washington

Citation:
AI-powered headphones offer group translation with voice cloning and 3D spatial audio (2025, May 10)
retrieved 10 May 2025
from https://techxplore.com/news/2025-05-ai-powered-headphones-group-voice.html

This document is subject to copyright. Apart from any fair dealing for the purpose of private study or research, no
part may be reproduced without the written permission. The content is provided for information purposes only.





Source link

Previous Post

Renewable energy: Closing financing gap in Global South requires multi-pronged approach – EnviroNews

Next Post

Galvion’s CORTEX Turns Combat Helmets Into Smart Mission Systems

Next Post
Galvion’s CORTEX Turns Combat Helmets Into Smart Mission Systems

Galvion’s CORTEX Turns Combat Helmets Into Smart Mission Systems

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

RECOMMENDED NEWS

Zimbabweans accuse Chinese investors of degrading environment

Zimbabweans accuse Chinese investors of degrading environment

8 months ago
President Ramkalawan Attends Durbar Marking 100th Anniversary of King Prempeh I’s Return from Seychelles Exile

President Ramkalawan Attends Durbar Marking 100th Anniversary of King Prempeh I’s Return from Seychelles Exile

6 months ago
UK police to trial new forensic footwear identification process

UK police to trial new forensic footwear identification process

1 year ago
Russia’s Tokenization Move: A Strategy for De-Dollarization?

Russia’s Tokenization Move: A Strategy for De-Dollarization?

6 months ago

POPULAR NEWS

  • Ghana to build three oil refineries, five petrochemical plants in energy sector overhaul

    Ghana to build three oil refineries, five petrochemical plants in energy sector overhaul

    0 shares
    Share 0 Tweet 0
  • When Will SHIB Reach $1? Here’s What ChatGPT Says

    0 shares
    Share 0 Tweet 0
  • Matthew Slater, son of Jackson State great, happy to see HBCUs back at the forefront

    0 shares
    Share 0 Tweet 0
  • Dolly Varden Focuses on Adding Ounces the Remainder of 2023

    0 shares
    Share 0 Tweet 0
  • US Dollar Might Fall To 96-97 Range in March 2024

    0 shares
    Share 0 Tweet 0
  • Privacy Policy
  • Contact

© 2023 LBNN - All rights reserved.

No Result
View All Result
  • Home
  • Business
  • Politics
  • Markets
  • Crypto
  • Economics
    • Manufacturing
    • Real Estate
    • Infrastructure
  • Finance
  • Energy
  • Creator Economy
  • Wealth Management
  • Taxes
  • Telecoms
  • Military & Defense
  • Careers
  • Technology
  • Artificial Intelligence
  • Investigative journalism
  • Art & Culture
  • Documentaries
  • Quizzes
    • Enneagram quiz
  • Newsletters
    • LBNN Newsletter
    • Divergent Capitalist

© 2023 LBNN - All rights reserved.