从工具替代到人机协作:深度推理大语言模型评阅医学生反思性写作作业的个案研究
From substitution to human-AI collaboration: a case study on assessment of reflective writing assignments of medical students using reasoning-capable large language models
摘要本研究对具有深度推理功能的大语言模型(reasoning-capable large language models,rcLLMs)评阅医学生反思性写作作业的可行性和效果进行了探索。在选择作业样本、建立评阅方法的基础上,用2种国内主流rcLLMs进行评阅,并采用定量与定性相结合的混合研究方法将评阅结果与教师评阅结果及评分标准相比较。研究发现2种rcLLMs和教师三者评分的一致性及两两之间评分的相关性较差,2种rcLLMs各自评分的稳定性也较差;rcLLMs生成的评语结构完整,但关注主题与评分标准不完全一致。本研究揭示出目前rcLLMs尚不能作为评阅医学生反思性写作作业的有效工具,讨论了相关影响因素,并强调了建立人机协作策略及反思AI评阅正当性的重要性。
更多相关知识
abstractsThis case study explored the feasibility and efficacy of employing reasoning-capable large language models (rcLLMs) to assess the reflective writing assignments of medical students. Following the selection of writing samples and the establishment of an assessment methodology, two prominent domestic rcLLMs were utilized to evaluate the assignments. A method combining quantitative and qualitative analyses was employed to compare the evaluations from rcLLMs with those from instructors in accordance with the established grading rubric. The findings revealed poor inter-rater reliability due to weak correlations among all three raters collectively and pairwise. Both rcLLMs demonstrated significant variability in their own scoring consistency. While the feedback comments generated by the rcLLMs were well-structured, their focuses exhibited a misalignment with the grading rubric. This study suggests that current rcLLMs are not yet suitable for use as effective tools for assessing the reflective writing assignments of medical students. The influencing factors were discussed, and the need to develop human-AI collaborative strategies, as well as to critically reflect on the legitimacy of employing rcLLMs in assessment, was underscored.
More相关知识
- 浏览1
- 被引0
- 下载0

相似文献
- 中文期刊
- 外文期刊
- 学位论文
- 会议论文


换一批



