Abstract:
Head-related transfer function (HRTF) plays an important role in immersive spatial audio playback. However, measuring or calculating HRTF for each person requires a lot of time and labor costs, which greatly limits the popularization and application of HRTF personalization. To address this problem, a personalized HRTF prediction method based on auditory fingerprint analysis is proposed. Auditory fingerprint refers to a personalized feature set extracted from individual head and torso measurement parameters and ear images. The ear image is used to replace the auricle measurement parameters, thereby avoiding the tedious manual measurement process. The geometric information contained in the ear image can effectively reflect the individual's ear morphological characteristics. Combined with the head and torso measurement parameters, the model's ability to model individual differences is enhanced. Through the training of the deep learning model, the system can extract features related to HRTF from these inputs and perform simultaneous prediction of amplitude and phase. Experimental results show that the proposed method has a log spectral distortion (LSD) of 4.33 dB and a phase root mean square error (RMSE) of 1.22 in the full frequency band, indicating that the proposed method can accurately predict personalized HRTF and has significant advantages over traditional methods. By constructing auditory fingerprints and combining deep learning models, the proposed method effectively avoids the data collection bottleneck in traditional methods and improves the efficiency and practicality of personalized HRTF prediction.