상세 보기
초록
Text-to-SQL is one of semantic parsing methods that converts natural language questions into SQL queries, and it aims to extract data from any relational database without knowledge of SQL query configuration. Although development of large amounts of datasets (WikiSQL, SPIDER) and development of pre-trained language models (BERT) contributed to the improvement of Text-to-SQL performance in English, language-specific dataset construction and model research have not been much progressed. Therefore, this study proposes a multilingual BERT-based Text-to-SQL methodology that converts the natural language question in Korean into SQL query for an English database. To this end, four strategies for translating Korean queries into English were explored, and their effectiveness was verified by applying each strategy to three text-to-SQL model structures. As a result of the experiment, it was confirmed that it showed a significant SQL generation performance even for Korean questions. The proposed methodology is meaningful in that it shows semantic inferences between database tables, column information, and questions composed of different languages are possible, and it is expected to support efficient database access by Korean users who lack proficiency in writing SQL queries.
키워드
- 제목
- 다국어 BERT를 활용한 한국어 자연어 질의의 SQL 변환
- 제목 (타언어)
- Text-to-SQL for Korean Language based on Multilingual BERT
- 저자
- 윤훈상; 허재혁; 김정섭; 강필성
- 발행일
- 2022
- 저널명
- 대한산업공학회지
- 권
- 48
- 호
- 1
- 페이지
- 91 ~ 104