多特征融合的英语口语考试自动评分系统的研究

李艳玲; 颜永红

doi:10.3724/SP.J.1146.2012.00172

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名

邮箱

手机号码

标题

留言内容

验证码

多特征融合的英语口语考试自动评分系统的研究

doi: 10.3724/SP.J.1146.2012.00172

李艳玲^1,,
颜永红¹

基金项目:

国家自然科学基金(10925419, 90920302, 10874203, 60875014, 61072124, 11074275, 11161140319)资助课题

计量
- 文章访问数: 2926
- HTML全文浏览量: 146
- PDF下载量: 874
- 被引次数: 0
出版历程
- 收稿日期: 2012-02-24
- 修回日期: 2012-06-05
- 刊出日期: 2012-09-19

Research for Automatic Short Answer Scoring in Spoken English Test Based on Multiple Features

LI Yan-Ling^1
,,
Yan Yong-Hong- ¹

摘要

摘要: 该文主要针对大规模英语口语考试自动评分系统的问答题型，采用多特征融合的方法进行评分。以语音识别文本作为研究对象，提取了3类特征进行评分。这3类特征分别是：相似度特征、句法特征和语音特征。总共9个特征从不同方面描述了考生回答与专家评分之间的关系。在相似度特征中，改进了Manhattan距离作为相似度。同时提出了基于编辑距离的关键词覆盖率的特征，充分考虑了识别文本中存在的单词变异现象，为给考生一个客观公平的分数提供依据。所有提取的特征利用多元线性回归模型进行融合，得到机器评分。实验结果表明，提取的特征对机器评分是十分有效的，并且在以考生为单位的系统评分性能达到了专家评分性能的98.4%。
- 自动语音识别 /
- 自动评分 /
- 特征选择 /
- 相似度 /
- 句法树
Abstract: This paper focuses on automatic scoring about ask-and-answer item in large scale of spoken English test. Three kinds of features are extracted to score based on the text from Automatic Speech Recognition (ASR). They are similarity features, parser features and features about speech. All of nine features describe the relation with human raters from different aspects. Among features of similarity measure, Manhattan distance is converted into similarity to improve the performance of scoring. Furthermore, keywords coverage rate based on edit distance is proposed to distinguish words variation in order to give students a more objective score. All of those features are put into multiple linear regression model to score. The experiment results show that performance of automatic scoring system based on speakers achieves 98.4% of human raters.
- Automatic Speech Recognition (ASR) /
- Automatic scoring /
- Feature selection /
- Similarity measure /
- Parser tree