Responsibilities
๋น์์ด๋ฏผ ์๋ ์์ฑ์ ํน์ฑ(๋ฐ์ ์ค์ฐจ, ๋ถ์ ํ์ฑ ๋ฑ)์ ๊ทน๋ณตํ๊ธฐ ์ํ ํ์ธ ํ๋ ๋ฐ์ดํฐ์ ์ ์ ๋ฐ ๋ชจ๋ธ ๋ฏธ์ธ ์กฐ์
Build fine-tuning datasets and optimize speech recognition models to improve performance on non-native children's speech, addressing pronunciation errors and speech variability.
GOP ์๊ณ ๋ฆฌ์ฆ ๊ธฐ๋ฐ์ ๋ฐ์ ํ๊ฐ ๋ฐ ์์ ๋ ๋ฒจ ์ ์ ์์คํ ์ค๊ณ ๋ฐ ๊ณ ๋ํ
Design, develop, and enhance pronunciation assessment systems and phoneme-level scoring based on the Goodness of Pronunciation (GOP) algorithm.
๊ตญ์ฑ ์ฐ๊ตฌ๊ธฐ๊ด(ETRI ๋ฑ)์ ์์ ๋จ์ ๋ฐ์ ๋ถ์ ๋ฐ ์์ ๋ฐํํ ์์ฑ์ธ์ ์์ง ์์ค์ฝ๋ ๋ถ์ ๋ฐ ๋ด์ฌํ
Analyze, adapt, and integrate phoneme-level pronunciation analysis and spontaneous speech recognition engine source code developed by national research institutes (e.g., ETRI).
Qualifications
์์ฑ ์ธ์(STT) ๋ฐ ์์ฐ์ด ์ฒ๋ฆฌ ํ์ดํ๋ผ์ธ ๊ตฌ์ถ ๊ฒฝํ์ด ์์ผ์ ๋ถ
Experience building speech recognition (STT) and natural language processing (NLP) pipelines.
์ํฅ ๋ชจ๋ธ ๋ฐ ์์ ๋จ์ ๋ถ์ ๊ธฐ์ ์ ๋ํ ๊น์ ์ดํด๊ฐ ์์ผ์ ๋ถ
Strong understanding of acoustic models and phoneme-level speech analysis techniques.
๊ธฐ์กด์ ํ์ต๋ ์์ฑ/NLP ์์ง์ ํน์ ๋๋ฉ์ธ(์๋ ๋ฐ์ ๋ฑ)์ ๋ง์ถฐ ํ์ธ ํ๋ํด ๋ณธ ๊ฒฝํ์ด ์์ผ์ ๋ถ
Experience fine-tuning pre-trained speech recognition or NLP models for domain-specific applications, such as children's pronunciation.
ETRI, ์คํ์์ค ๋ฑ ์ธ๋ถ ๊ธฐ๊ด์ AI API ๋๋ ์์ค์ฝ๋๋ฅผ ์ด์ ๋ฐ์ ์๋น์ค์ ํฌํ ํด ๋ณธ ๊ฒฝํ์ด ์์ผ์ ๋ถ
Experience integrating AI APIs or source code from external organizations (e.g., ETRI or open-source projects) into production services.
์์ฑ ์ธ์ ๋ชจ๋ธ์ ๋ชจ๋ฐ์ผ ํ๋ซํผ ์ด์ ๊ฒฝํ์ด ์์ผ์ ๋ถ
Experience deploying and optimizing speech recognition models on mobile platforms.
Tech Stack
Python, PyTorch / TensorFlow, Hugging Face, Librosa, Audio Processing Tools
๋๊ธ