박용화 교수 연구팀, 회전하는 스피커 어레이와 딥러닝을 이용한 개인 맞춤형 입체음향 전달함수 측정 기술 개발
2026.09.03
admin
KAIST 기계공학과 박용화 교수 연구팀이 연속 회전하는 스피커 어레이와 심층신경망(DNN)을 결합해 개인 맞춤형 머리전달임펄스응답(HRIR)을 단 5초의 측정만으로 정밀하게 추정하는 기술을 개발하고, 그 결과를 계측 분야 저명 국제 학술지인 IEEE Transactions on Instrumentation and Measurement (Impact Factor: 7.0, 1저자 고병윤 박사)에 게재하였다.

VR·AR과 메타버스 환경에서 자연스러운 공간 음향을 구현하려면, 소리가 사람의 머리·귓바퀴·몸통을 거쳐 귀에 도달하는 과정을 기술한 HRIR이 필요하다. 그런데 HRIR은 머리와 귀의 형상에 따라 개인차가 매우 커서, 타인의 평균 데이터를 그대로 사용하면 앞뒤 방향을 혼동하거나 소리가 머리 안쪽에서 들리는 현상이 발생한다. 이는 단순한 몰입감 저하를 넘어, 시각장애인용 음향 내비게이션이나 난청 환자의 청각 재활 훈련에서는 실질적인 문제로 이어진다. 문제는 개인 HRIR을 얻으려면 스피커를 방위각마다 옮겨가며 반복 측정해야 해 시간이 오래 걸리고, 그동안 피험자가 움직이지 않고 있어야 한다는 점이다.
측정 시간을 줄이기 위해 스피커 어레이를 계속 회전시키면서 측정하는 방식이 제안되어 왔지만, 회전 속도가 빨라질수록 HRIR이 급격히 변해 기존의 수식 기반 적응 필터(NLMS, 칼만 필터 등)로는 그 변화를 따라가지 못하는 한계가 있었다. 연구팀은 이 문제를 해결하기 위해 GRU 구조와 완전연결 신경망을 결합한 DNN 모델을 설계하고, 시퀀스-투-시퀀스 학습으로 시간에 따라 변하는 HRIR을 연속적으로 추정하는 방법을 제안했다. 특히 스피커 구동 신호의 세기를 이용해 오차 기울기의 크기를 스스로 조절하는 학습형 정규화(learnable normalization) 기법과, 전체 시퀀스를 모두 처리한 뒤에 학습을 수행하는 전체 시퀀스 갱신·최적화 기법을 도입해, 별도의 학습 데이터셋 없이도 안정적인 추정이 가능하도록 했다. 그 결과 무향실 실험에서 기존 기법 대비 약 6 dB 향상된 정확도를 달성했으며, 정상 청력 성인 10명을 대상으로 한 가상환경 음원 정위 실험에서는 범용 데이터 대비 수평면 약 16°, 정중면 약 39°의 정위 오차 감소를 보였다. 측정에 걸린 시간은 단 5초였다.


박용화 교수는 "짧은 측정 시간과 높은 정확도를 동시에 확보함으로써, 개인 맞춤형 입체음향을 실험실 밖에서도 구현할 수 있는 가능성을 열었다"며, "VR·AR 콘텐츠뿐 아니라 시각장애인 음향 내비게이션과 난청 환자의 청각 훈련 등 사회적 활용도가 높은 분야에 기여할 수 있을 것"이라고 밝혔다.
이 논문은 방위사업청과 산업통상자원부의 재원으로 민군협력진흥원의 지원을 받아 수행된 연구임(23-SN-CV-04). 본 연구는 과학기술정보통신부에서 지원하는 과학기술원 InnoCORE 사업에 의해 수행되었습니다(N10250154).
자세한 정보와 전체 논문은 IEEE Transactions on Instrumentation and Measurement 또는 IEEE Xplore를 참조하세요.
(https://doi.org/10.1109/TIM.2025.3644532)

Professor Yong-Hwa Park's research group developed a deep learning method that identifies individualized head-related impulse responses in a 5-second measurement using a continuously rotating speaker array
The research team led by Professor Yong-Hwa Park of the Department of Mechanical Engineering at KAIST has published a research paper in IEEE Transactions on Instrumentation and Measurement (Impact Factor: 7.0, First Author: Dr. Byeong-Yun Ko), a renowned international journals in the field of instrumentation and measurement. The work enables accurate identification of individualized head-related impulse responses (HRIRs) from a measurement lasting only five seconds, using a continuously rotating speaker array combined with a deep neural network (DNN).
Rendering natural spatial audio in VR, AR, and metaverse environments requires HRIRs, which describe how sound from a given direction reaches the ear after interacting with the listener's head, pinna, and torso. Because HRIRs are highly individual, using generic, nonindividualized data causes perceptual errors such as front–back confusion and in-head localization. Beyond degraded immersion, these errors have practical consequences in audio navigation for the visually impaired and auditory localization training for patients with hearing loss. Conventional static measurement, however, requires repositioning the speaker array for every azimuth angle, which takes considerable time and demands that the subject remain motionless throughout.
Dynamic approaches that rotate the speaker array continuously have been proposed to shorten measurement time, but their accuracy degrades sharply at high rotational speeds, since analytical adaptive filters such as NLMS and the Kalman filter cannot track the rapid HRIR transitions. To address this, the team designed a DNN that integrates gated recurrent unit (GRU) structures with fully connected networks and performs sequence-to-sequence learning to update the HRIR vector over time. A learnable normalization based on the power of the speaker excitation signal adaptively stabilizes the scale of the instantaneous square error gradient and the update rate, and a whole-sequence updating and optimization scheme enables the model to be optimized without requiring any training dataset, while preventing overfitting. As a result, the method achieved approximately 6 dB higher accuracy than conventional approaches in anechoic chamber experiments, and in a virtual-environment localization test with ten normal-hearing subjects, the individualized HRIRs reduced localization error by about 16° on the horizontal plane and 39° on the median plane compared with generic data — from a measurement that took only five seconds.
Professor Yong-Hwa Park stated, "By achieving short measurement time and high accuracy at the same time, this work opens the possibility of realizing individualized spatial audio outside the laboratory." He added, "Beyond VR and AR content, the method can contribute to socially valuable applications such as audio navigation for the visually impaired and auditory training for patients with hearing loss.“
This work was supported in part by the Institute of Civil Military Technology Cooperation through the Defense Acquisition Program Administration and the Ministry of Trade, Industry and Energy of the Korean Government under Grant 23-SN-CV-04, and in part by the InnoCORE Program of the Ministry of Science and ICT under Grant N10250154.
For detailed information and the full paper, please refer to IEEE Transactions on Instrumentation and Measurement or IEEE Xplore (https://doi.org/10.1109/TIM.2025.3644532).