Přístupnostní navigace
E-application
Search Search Close
Master's Thesis
Author of thesis: Ing. Vojtěch Rezek
Acad. year: 2025/2026
Supervisor: Ing. Ondřej Mokrý, Ph.D.
Reviewer: Ing. Michal Švento
This thesis addresses the problem of reconstructing degraded audio signals using deep learning, with a primary focus on how different time-frequency signal representations affect reconstruction quality. Three neural network models were implemented and compared, each employing a distinct representation: Short-Time Fourier Transform (STFT), Wavelet Transform, and Constant-Q Transform (CQT). The models were trained and evaluated on the real-world IRMAS dataset of musical recordings, with two types of degradation simulated: saturation and additive noise. Results were assessed using both objective and perceptual metrics (SNR, ODG, PSM). The experiments demonstrated that the choice of signal representation has a substantial and measurable impact on reconstruction quality -- differences between configurations exceed 6~dB in SNR and span an entire category on the perceptual ODG scale. The best perceptual quality for noise reduction was achieved by a wavelet configuration with ODG~$-0.382$, corresponding to the \textit{just noticeable degradation} category on the PEMO-Q scale. The thesis also contributes a custom CQT processing pipeline for MATLAB and a block segmentation strategy enabling efficient training on standard CPU hardware without GPU acceleration.
signal reconstruction, deep learning, neural networks, time-frequency analysis, STFT, wavelet transform, CQT, audio processing, audio restoration, psychoacoustics
Date of defence
11.06.2026
Result of the defence
Defended (thesis was successfully defended)
Grading
E
Process of defence
Student prezentoval výsledky své práce a komise byla seznámena s posudky. Student obhájil diplomovou práci s výhradami a odpověděl na otázky členů komise a oponenta. Otázky: 1) V texte popisujete časovú náročnosť pri trénovaní, no nepopisujete aký je čas na priechod sieťou po natrénovaní v tzv. inferencii, prosím doplňte analýzu predložených sietí v inferencii. 2) V kapitole 5. 2. predstavujete jednoduchý model s STFT, v ďalších kapitolách „zložitejšie“ modely, ktoré sú ale veľmi podobných rozmerov. Vysvetlite, prečo ste použili tak malé architektúry sietí. 3) Uveďte počet učiteľných parametrov vašich sietí a vysvetlite, či sú všetky vrstvy učiteľné. 4) Vysvětlete pojem ERBlet.
Language of thesis
Czech
Faculty
Fakulta elektrotechniky a komunikačních technologií
Department
Department of Telecommunications
Study programme
Audio Engineering (MPC-AUD)
Specialization
Audio Production and Recording (AUDM-ZVUK)
Composition of Committee
Ing. Jaromír Mačák, Ph.D. (člen) Doc.Ing.MgA. Ondřej Urban, Ph.D. (předseda) doc. Ing. Jiří Schimmel, Ph.D. (místopředseda) RNDr. Lubor Přikryl (člen) Ing. Ondřej Mokrý, Ph.D. (člen)
Supervisor’s reportIng. Ondřej Mokrý, Ph.D.
Grade proposed by supervisor: E
Reviewer’s reportIng. Michal Švento
Grade proposed by reviewer: F
Responsibility: Mgr. et Mgr. Hana Odstrčilová