상세 보기
Designing Speech-to-Speech Pipelines for Effective English Conversation Learning: A Modular VAD–STT–LLM–TTS Perspective
초록
This paper reconceptualizes Speech-to-Speech (STS) systems for English conversation education through a modular pipeline lens. We argue that educational effectiveness depends on intentionally designing the VAD–STT–LLM–TTS pipeline as an integrated pedagogical system rather than treating STS as a black box. We summarize the functional roles of each component and examine how they shape key dimensions of conversation learning, including fluency, accuracy, interactional competence, affective support, and pronunciation/listening development. We also highlight cross-module dynamics such as error propagation and compensation and discuss how end-to-end latency influences learners’ sense of real-time interaction. Based on these insights, we propose practical design principles for classroom and self-study contexts, including task- and proficiency-sensitive VAD/STT settings, level-adaptive LLM prompting and feedback strategies, and activity designs leveraging controllable TTS for listen-and-repeat, shadowing, and role-play. We conclude with implications for future empirical validation and learner- and domain-specific optimization.
키워드
- 제목
- Designing Speech-to-Speech Pipelines for Effective English Conversation Learning: A Modular VAD–STT–LLM–TTS Perspective
- 저자
- 남호성
- 발행일
- 2026-02
- 유형
- Y
- 저널명
- 외국어교육연구
- 권
- 40
- 호
- 1
- 페이지
- 19 ~ 48