Last updated: 2026-01-31
- LLM-based conference manager: Upload meeting audio → STT transcription → Summary & Q&A chatbot
- Collaborated with 2 developers (AI, Frontend/Backend)
- My contribution
- Data collection (for fine tuning Whisper STT model)
- Collected ~1000 hours of Korean YouTube audio with manual subtitles
- 10 different domains for diverse vocabulary and audio quality
- Audio preprocessing
- Extracted audio from YouTube, split by subtitle timestamps
- Merged segments into 30-second chunks
- Converted sampling rate & channels to match Whisper input format
- Result: Dataset size reduced from 707GB → 118GB, upload time 8hr → 1hr 42min
- Text preprocessing & CER filtering
- Compared raw subtitles vs Whisper transcripts
- Filtered data with CER > 30 to remove low-quality samples
- Prompt engineering, backend process/algorithm/api, RAG...
- Data collection (for fine tuning Whisper STT model)