Source-linked AI summary
Sensing Bone-Conducted Speech with Earbuds
Christoph Weyer, Peter Jax
TL;DR
Clear OV capture with earbuds is difficult in acoustic noise, while the bandwidth and spatial characteristics of OV-induced earbud vibrations had not been analyzed in detail. This study measures those characteristics with two earbud models and finds low-pass vibrations concentrated below 400 Hz, with consistent in-and-out motion that supports single-axis sensing for high-power components.
Problem
OV capture is challenging in noisy environments, and the bandwidth and spatial characteristics of OV-induced earbud vibrations were not analyzed in detail despite their relevance to sensor choice and placement.
Method
The study measures OV-induced accelerations with three-axis accelerometers on two earbud models, using reference speech and head-orientation measurements to analyze spectral and spatial characteristics.
Results
−93 dB per decade above 400 Hz; vibrations are mainly in and out of the ear canal entrance, and favorable single-axis sensing yields less than 1.5 dB mean attenuation below 400 Hz.
Takeaways & Limitations
The high-power vibration components below 400 Hz are the most relevant for many applications, while frequencies above 1 kHz require comparatively low-noise-floor sensors.
Takeaways & Limitations
The analysis cannot determine whether lower-noise accelerometers would recover more signal than single-axis projection losses because the three-axis accelerometer’s noise limits the analysis.
Abstract
from arXiv · showhide
Clear capture of the wearer's own voice (OV) is essential when using earbuds for mobile communication. However, OV capture remains challenging in noisy environments. Bone-conducted (BC) speech, which can be sensed as vibrations of the earbud housing, can be used to improve OV capture. However, neither bandwidth nor spatial characteristics of OV-induced earbud vibrations have been analyzed in detail, despite both characteristics being relevant, e.g., for sensor choice and placement. This study investigates both characteristics, based on measurements with two earbud models. Spectrally, results indicate that OV-induced earbud vibrations exhibit a low-pass characteristic, with a steep roll-off of -93 dB per decade above 400 Hz. Thus, sensors with comparatively low noise floors are required to sense the vibrations above \SI{1}{\kilo\hertz}. Spatially, results indicate that the earbuds mainly vibrate in and out of the ear canal entrance, with high consistency between subjects and fits. Simulations confirm that this enables capture of the high-power vibrations below 400 Hz by a single-axis sensor with less than 1.5 dB mean attenuation.
1 Introduction
The study addresses the limited analysis of OV-induced earbud vibrations, focusing on their spectral bandwidth and spatial orientation because both affect sensor selection and placement.
- Motivation: Bone-conduction microphones capture the wearer’s own voice as earbud or body vibrations and are robust against acoustic noise and wind.The paper uses “bone-conduction” for consistency with existing literature, while noting that “body-conduction” is more appropriate here.
- Research gap: Earbud pickup has become increasingly common for speech applications, but the fundamental characteristics of OV-induced earbud vibrations remain insufficiently analyzed.Existing work has used earbuds for speech enhancement, recognition, voice activity detection, wind-noise reduction, and occlusion-effect reduction.
- Practical relevance: Spectral bandwidth and spatial orientation determine accelerometer bandwidth, axis count, and mounting orientation, with trade-offs involving cost, processing complexity, and SNR.These characteristics are therefore relevant to earbud and sensing-component design.
- Contribution: The study extends prior pilot evidence using two earbud models, more participants, joint spectral-and-SNR analysis, head-related orientation estimation, and single-axis attenuation simulations.The prior pilot study involved five models but relatively few test subjects.
- Paper scope: The paper analyzes experimental setup, orientation estimation, spectral characteristics, spatial characteristics, and the effectiveness of single-axis sensing.These analyses are presented across Sections 2–7 before the conclusion.
2 Experiment Setup
The experiment measures OV-induced earbud accelerations, mouth-adjacent reference sound, and head orientation across subjects, earbud models, and realistic refits.
- Recorded signals: A headset microphone records mouth-adjacent sound pressure as a reference, while a head tracker records head orientation.Signals are routed through an audio codec and sound card to a computer.
- Signals and coordinate systems: The accelerometer records vibration along the three axes of its local coordinate system A, while rotation matrices relate accelerometer and head-tracker frames to world coordinates.The headset microphone signal s(k) represents sound pressure near the mouth.
- Hardware: The study equips Anker P3i and A20i earbuds with three-axis accelerometers to measure OV-induced housing accelerations.The P3i has a stem and silicone securing wing, while the knob-like A20i lacks a wing.
- Participants and fitting: Measurements include 17 subjects, two earbud pairs, and one refit for each subject–earbud combination.Participants were instructed to achieve a tight ear-canal seal and stable fit.
- Procedure: Each fit includes calibration, text, and additional recordings, with accelerometer and microphone signals jointly sampled at 48 kHz.Head orientation is captured at 2 Hz throughout the recordings.
3 Estimation of Earbud Orientation
The paper estimates earbud orientation relative to the wearer’s head, then expresses measured accelerations in a common head-aligned coordinate system for cross-subject and cross-earbud comparison.
- Purpose: The method transforms accelerometer measurements from each earbud’s local coordinate system into a fixed coordinate system aligned with the head.This enables visualization and comparison despite changes in earbud fit and orientation.
- Orientation estimation: Earbud orientation is estimated from gravity vectors recorded in accelerometer coordinates during two head positions.The method assumes known world-frame gravity, a constant earbud-to-head orientation, and a known head rotation between recordings.
- Calibration: During calibration, the accelerometer orientation in world coordinates is inferred from the mapped gravity vector and the known world-frame gravity direction.The calibration orientation is represented by Wcalib_A.
- Tilt recording: A forward tilt provides a second gravity observation, with the tilt orientation expressed through the head rotation ΔR applied to the calibration orientation.The paired observations form the equation system used to estimate the earbud orientation.
4 Typical Spectra of Earbud Vibration
The analysis estimates OV-induced earbud vibration spectra and their SNR limits using PSD/ASD measurements across axes and accelerometers. Vibrations show a strong low-pass characteristic: power is concentrated below 400 Hz, while sensing above 1 kHz requires substantially lower sensor noise floors.
- Methodology and Procedure: The study estimates OV and noise PSDs from text and silent calibration recordings, then computes frequency-dependent SNR for each accelerometer axis.The OV PSD is obtained by subtracting the calibration noise PSD from the text-recording PSD; negative estimates are set to zero.
- Results: Around 530 µg/√Hz, the mean OV ASD remains roughly constant between 100 Hz and 400 Hz, varying across axes, earbuds, and voice loudness.Speech harmonics can produce low ASD values between harmonics, especially below the fundamental frequency.
- Results: −93 dB per decade, the OV ASD rolls off steeply above 400 Hz, reflecting reduced speech excitation and damping during transmission to the earbud.Above 1 kHz, estimation error increases as the OV signal approaches the noise floor, making the estimated ASD an upper bound on the true value.
- Results: 50 µg/√Hz between 100 Hz and 400 Hz versus 0.5 µg/√Hz between 1 kHz and 2 kHz are the approximate noise floors needed to capture average ASD components at 20 dB.The steep roll-off therefore makes high-frequency components substantially more demanding to sense.
- Results: Above 1 kHz, accelerometers A and B have mean-OV SNRs below 0 dB, whereas accelerometer C reaches around 15 dB.The results suggest that sensing components up to 2 kHz likely requires a noise floor at least on par with accelerometer C, although practical feasibility is not conclusive.
- Results: High-power components are concentrated between 100 Hz and 400 Hz, while the SNR limits the bandwidth in which later analyses remain valid to below 1 kHz.Accelerometer A shows approximately flat self-noise, while B and C have lower noise ASDs that decrease toward higher frequencies.
5 Bandwidth for Different Earbuds
OV-induced earbud vibrations show a consistent low-pass transmission characteristic across both tested earbud models, with model-dependent cutoff frequencies and similar behavior across sides.
- Measurement and metric: ||Hsa(f)|| measures total earbud acceleration relative to headset-microphone sound pressure, quantifying the relation between bone- and air-conducted speech.The transfer function is computed from headset microphone to accelerometer axes and evaluated using its norm.
- Observed bandwidth: Both earbud models exhibit a similar low-pass characteristic with high consistency between left and right sides.Variation around the mean is limited above 200 Hz, while outliers below 200 Hz reflect subjects’ fundamental voice frequencies.
- Observed bandwidth: The cutoff begins around 280 Hz for the A20i and around 360 Hz for the P3i.After the cutoff, the mean transfer-function norms roll off by around −58 dB/decade.
- Interpretation: Components above 400 Hz are attenuated by transmission through tissue and the earbud, contributing to the observed low-pass behavior.The cutoff differences are attributed as likely resulting from earbud mass and coupling to the concha.
6 Spatial Characteristics of Vibration
Earbud vibrations are predominantly oriented along one direction, with modest variation across subjects and fits, and are directed out of the ear canal toward the sides of the head.
- Implication for sensing: The consistent dominant direction supports evaluating single-axis sensing as a way to capture most vibration power with reduced sensor-axis complexity.The paper uses spatial consistency as the basis for assessing single-axis capture effectiveness.
- 6.2 Directionality: PCA evaluates whether each three-axis accelerometer recording has a dominant spatial vibration direction across subjects, earbuds, fits, and sides.Signals are bandpass-filtered from 100 Hz to 1.5 kHz before PCA, and each principal component’s explained variance is computed.
- 6.2 Directionality: 89% of P3i variance is explained by the first principal component, compared with 79% for the A20i.For the P3i, the second and third components explain 9% and 2% on average; the A20i is less consistently one-directional.
- 6.2 Directionality: For some A20i recordings, the second principal component explains up to 44% of total variance, indicating more equal acceleration across directions in a plane.This second direction is especially relevant for some A20i signals.
- 6.3 Variance of Main Direction: Mean angular deviation is 6.5° for the P3i and 14.7° for the A20i, with maximum deviations of 22.2° and 38.9°, respectively.All maximum deviations remain within a ±45° cone around the mean direction.
- 6.3 Variance of Main Direction: In a head-aligned coordinate system, the main acceleration direction clusters by model, is largely symmetrical between sides, and points toward the sides of the head.The P3i trends downward and the A20i backward; the identified direction mainly reflects high-power components below 400 Hz.
7 Simulation of Single-Axis Accelerometer
The study simulates projecting three-axis earbud acceleration onto candidate single-axis directions and evaluates frequency-dependent attenuation. Favorably chosen directions capture high-power components effectively, although performance varies by earbud model and frequency.
- Simulation approach: A single-axis sensor is simulated by projecting the three-axis acceleration onto unit direction d and assessing the resulting attenuation per frequency.The power spectral gain measures the fraction of total vibration power captured along the selected direction.
- P3i results: For the P3i, the favorable direction P3i ap1 yields 0.7 dB mean attenuation from 100 Hz to 750 Hz, versus 14 dB for P3i ap3.The favorable direction also has relatively low variance, except toward frequencies below 200 Hz.
- A20i results: For the A20i, favorable direction A20i ap1 yields 2.4 dB mean attenuation from 100 Hz to 750 Hz, versus 9.7 dB for A20i ap3.Attenuation for the favorable direction varies substantially, reaching up to 13 dB around 600 Hz for some recordings.
- High-power region: Between 100 Hz and 400 Hz, favorable directions yield 0.7 dB and 1.5 dB mean attenuation for the P3i and A20i, respectively.Unfavorable directions yield 15 dB for the P3i and 11 dB for the A20i in the same high-power region.
- Model dependence: Single-axis capture works across a larger frequency range for the P3i, whereas the A20i may require three axes for accurate capture of higher frequencies.For the A20i, favorable-direction attenuation appears around 600 Hz, indicating more complex spatial behavior.
- Higher-frequency attenuation: Adding a second direction reduces attenuation around 600 Hz more effectively for many recordings when the second principal component is used.The two-degree-of-freedom PSG measures acceleration power within the plane spanned by two orthogonal directions.
- Mounting-direction tolerance: The normalized mean PSG remains within −3 dB for angular distances up to ±45° for the P3i and −2.7 dB for the A20i.Large angular distances, corresponding to near-perpendicular sensing, produce average attenuations of up to 12 dB.
8 CONCLUSION
The paper analyzes spectral and spatial characteristics of own-voice-induced earbud vibrations using two earbud models. It finds a pronounced low-pass spectrum and largely consistent vibration direction, supporting low-attenuation single-axis sensing of high-power components.
- Scope: The study analyzes spectral and spatial characteristics of earbud vibrations induced by the wearer’s own voice in two earbud models.The vibrations are evaluated as accelerations of the earbud housing.
- Spectral characteristics: OV-induced earbud vibrations have high-power components below 400 Hz and an approximate −93 dB-per-decade roll-off toward higher frequencies.The exact cutoff varies with earbud model, likely because of differences in coupling to the concha.
- Sensor implications: Sensing frequencies above 1 kHz requires comparatively low-noise-floor accelerometers, while multiple commercial accelerometers can sense high-power components above 20 dB average SNR.The high-power components below 400 Hz are therefore the most relevant for many practical applications.
- Spatial characteristics: High-power vibrations mainly accelerate the earbud in and out of the ear-canal entrance, with consistent direction across subjects, fits, and both earbud models.This consistency supports using a single-axis accelerometer to capture the high-power vibration components.
- Single-axis sensing: Simulations show that a favorably mounted single-axis sensor captures high-power vibrations below 400 Hz with less than 1.5 dB mean attenuation.The result links the observed spatial consistency to effective single-axis capture.