Designing Speech-to-Speech Pipelines for Effective English Conversation Learning: A Modular VAD–STT–LLM–TTS Perspective

초록

This paper reconceptualizes Speech-to-Speech (STS) systems for English conversation education through a modular pipeline lens. We argue that educational effectiveness depends on intentionally designing the VAD–STT–LLM–TTS pipeline as an integrated pedagogical system rather than treating STS as a black box. We summarize the functional roles of each component and examine how they shape key dimensions of conversation learning, including fluency, accuracy, interactional competence, affective support, and pronunciation/listening development. We also highlight cross-module dynamics such as error propagation and compensation and discuss how end-to-end latency influences learners’ sense of real-time interaction. Based on these insights, we propose practical design principles for classroom and self-study contexts, including task- and proficiency-sensitive VAD/STT settings, level-adaptive LLM prompting and feedback strategies, and activity designs leveraging controllable TTS for listen-and-repeat, shadowing, and role-play. We conclude with implications for future empirical validation and learner- and domain-specific optimization.

키워드

pipeline designVADSTTLLMTTSpedagogical pipeline tutoring음성-대-음성 변환영어회화 학습파이프라인 설계음성 활동 감지음성-텍스트 변환대규모 언어 모델텍스트-음성 변환교육적 파이프라인 관점의 튜터링speech-to-speechEnglish conversation learning
제목
Designing Speech-to-Speech Pipelines for Effective English Conversation Learning: A Modular VAD–STT–LLM–TTS Perspective
저자
남호성
DOI
10.16933/sfle.2026.40.1.19
발행일
2026-02
유형
Y
저널명
외국어교육연구
40
1
페이지
19 ~ 48