Make sure your voice truly gets through — stress-free online meetings for everyone.
…But what does it mean when online meetings are "hard to hear"?
Academic research on evaluating and improving speech quality in online meetings
Motivation
In online meetings, sometimes speakers’ voices are buried in noise, hard to hear, or simply difficult to concentrate on.
Many meetings also include power hierarchy: even if a manager or other senior person has poor audio quality and is hard to understand, it can be difficult for others to point this out.
Even when the meeting itself appears to proceed without major issues, you may later find that AI transcription or minutes contain many errors, making it time-consuming to confirm what was actually discussed.
Situations like these are commonplace, and they strongly affect how comfortable online meetings feel, how easy they are to follow, how lively the discussion becomes, and ultimately the productivity of the organization.
To check how your own voice sounds to others, you have to record yourself using the meeting tool’s audio test feature and listen back. In a busy work environment, however, it is hard to make time for this before every meeting, and the judgment of “easy to hear or not” relies entirely on personal, subjective impressions.
In addition, even if you feel the audio is hard to hear, it is not easy for a typical user to identify the underlying cause.
We believe that, if there were a service that could automatically analyze your audio quality and suggest improvements based on the specific causes of degradation, online meetings could become more comfortable and efficient.
Likewise, in AI transcription and automated minutes services, if we could know in advance whether the audio quality is suitable for accurate recognition, we could further improve the productivity of knowledge work.
Problem statement
In this research, we start from the hypothesis:
“If AI can correctly recognize a speech signal, that speech is also easy for humans to understand.”
Based on this assumption, we analyze how various factors that degrade audio affect the accuracy of AI transcription.
We then train AI models to learn the relationship between each degradation factor and transcription accuracy, exploring the possibility of objectively evaluating the quality of online meeting audio.

Examples of degraded audio
by 40% packet loss
(no packet loss)
Are the tendencies of mishearing different between humans and AI?
Research methodology

We use simulation to intentionally generate degraded speech and train AI models to learn the patterns of degradation.
Using these models, we:
- Analyze the causes and severity of audio degradation
- Propose specific improvement strategies based on those causes
- Apply the learned models to real meeting audio to automatically diagnose issues and provide actionable feedback
This research project is based on the hypothesis: "If AI can correctly recognize a speech signal, that speech is also easy for humans to understand".
To verify this, we designed a subjective listening experiment that observes how human listeners actually perceive degraded speech.
Among the many factors that affect speech quality, our first step focuses on packet loss in IP networks. We simulated packet loss in Japanese speech signals and conducted a transcription task using these degraded audio samples. By analyzing what participants wrote down and how they did so, we examine how packet loss affects:
- Human perception (clarity, intelligibility, understanding)
- Human behavior (number of replays, time required to complete the transcription, etc.)
We then compare the human transcription results with those of an automatic speech recognition (ASR) system and analyze how error patterns differ depending on acoustic and linguistic factors. This allows us to quantitatively understand the differences between how humans and AI “listen”, and to explore:
- More advanced automatic speech quality assessment methods
- Human-centered voice interface designs that are grounded in human perception and are genuinely easy to use
Conference & Publications
6th Joint Meeting Acoustical Society of America and Acoustical Society of Japan
Honolulu, Hawaii, USA, December 1-5, 2025
Presentation Title: Evaluating speech quality for automatic transcription in videoconferencing
(Press Release) Acoustics Lay Language Papers (Acoustical Society of America)
Article title: How Online Meetings Change Your Voice—and How We Measure It
Article URL: https://acoustics.org/how-online-meetings-change-your-voice-and-how-we-measure-it/
Acoustical Society of Japan, 2026 Spring Meeting
Tokyo, Japan, March 17-19, 2026
Poster presentation
Presentation Title: The Effects of Packet Loss in IP Communication on the Transcription Accuracy of Japanese Speech
(first day of the conference: 3/17)
Acoustical Society of Japan, 2026 Autumn Meeting
Kanazawa city, Ishikawa pref., Japan, September 8-10, 2026
Poster presentation
Presentation Title: Comparison of Human and Automatic Speech Recognition Responses to Packet-Loss-Degraded Speech
(Second day of the conference: 9/9)
Presentation Title: Automatic Speech Recognition–Based Speech Quality Assessment
(Third day of the conference: 9/10)