• 中国期刊全文数据库
  • 中国学术期刊综合评价数据库
  • 中国科技论文与引文数据库
  • 中国核心期刊(遴选)数据库
焦顺, 王玫. 基于听觉指纹分析的头相关传输函数个性化J. 桂林电子科技大学学报, xxxx, x(x): 1-8. DOI: 10.16725/j.1673-808X.202558
引用本文: 焦顺, 王玫. 基于听觉指纹分析的头相关传输函数个性化J. 桂林电子科技大学学报, xxxx, x(x): 1-8. DOI: 10.16725/j.1673-808X.202558
JIAO Shun, WANG Mei. Personalized head related transfer function based on auditory fingerprint analysisJ. Journal of Guilin University of Electronic Technology, xxxx, x(x): 1-8. DOI: 10.16725/j.1673-808X.202558
Citation: JIAO Shun, WANG Mei. Personalized head related transfer function based on auditory fingerprint analysisJ. Journal of Guilin University of Electronic Technology, xxxx, x(x): 1-8. DOI: 10.16725/j.1673-808X.202558

基于听觉指纹分析的头相关传输函数个性化

Personalized head related transfer function based on auditory fingerprint analysis

  • 摘要: 头相关传输函数(HRTF)在沉浸式空间音频重放中扮演着重要的角色。然而为每个人测量或计算HRTF需要消耗大量的时间和人力成本,这极大地限制了HRTF个性化的普及与应用。针对这一问题,提出了一种基于听觉指纹分析的HRTF个性化方法,听觉指纹是指通过个体的头部、躯干测量参数以及耳朵图像等信息提取的个性化特征集合,通过使用耳朵图像来代替耳廓测量参数,从而避免了繁琐的人工测量过程,且耳朵图像中包含的几何信息可以有效反映个体的耳部形态特征,结合头部和躯干的测量参数,从而增强了模型对个体差异的建模能力,通过深度学习模型的训练,系统能够从这些输入中提取出与HRTF相关的特征,并进行幅度和相位的同时预测。实验结果表明,所提出的方法在全频带范围内的对数谱失真为4.33 dB,相位均方根误差为1.22,表明该方法能够准确预测个性化HRTF,且相比传统方法具有显著的优势。通过构建听觉指纹并结合深度学习模型,所提方法有效避免了传统方法中的数据采集瓶颈,提高了个性化HRTF预测的效率和实用性。

     

    Abstract: Head-related transfer function (HRTF) plays an important role in immersive spatial audio playback. However, measuring or calculating HRTF for each person requires a lot of time and labor costs, which greatly limits the popularization and application of HRTF personalization. To address this problem, a personalized HRTF prediction method based on auditory fingerprint analysis is proposed. Auditory fingerprint refers to a personalized feature set extracted from individual head and torso measurement parameters and ear images. The ear image is used to replace the auricle measurement parameters, thereby avoiding the tedious manual measurement process. The geometric information contained in the ear image can effectively reflect the individual's ear morphological characteristics. Combined with the head and torso measurement parameters, the model's ability to model individual differences is enhanced. Through the training of the deep learning model, the system can extract features related to HRTF from these inputs and perform simultaneous prediction of amplitude and phase. Experimental results show that the proposed method has a log spectral distortion (LSD) of 4.33 dB and a phase root mean square error (RMSE) of 1.22 in the full frequency band, indicating that the proposed method can accurately predict personalized HRTF and has significant advantages over traditional methods. By constructing auditory fingerprints and combining deep learning models, the proposed method effectively avoids the data collection bottleneck in traditional methods and improves the efficiency and practicality of personalized HRTF prediction.

     

/

返回文章
返回