To introduce the second month of our mini series in the influence of generative artificial intelligence (GenAI) in languages, from how we speak and write, to how we learn languages, I want to introduce you to some literature concerning the use of chatbots employing large language models (LLMs) in language learning.
In the ecosystem of language learning applications, there have been several notable shifts thought to streamline or personalise the user experience in their target languages. Beyond the traditional classroom setting and coursebooks, numerous language learning applications have been popping up in App Stores, increasing the accessibility of learning a foreign language to basically anyone who is interested. Algorithms meant to tailor user exercises have since developed, using various methods such as spaced repetition to improve user retention of newly learned words.
With the rise of generative artificial intelligence (GenAI), we are starting to see another shift in how language learning applications are designed, though reactions are largely mixed. GenAI has given chatbots a bump in their capabilities through developments in natural language processing and large language models, which have grown increasingly capable of communicating with human users. This has undoubtedly drawn some scrutiny and questions, such as if these GenAI-powered chatbots are indeed effective in language education.
While a relatively nascent method, there are several systematic reviews and meta-analyses that have been conducted to compile the existing evidence of the impact of chatbot use or integration in foreign language learning, with some focusing on efficacies, and others generally touching on the purported benefits and challenges of such implementations. Today, we will look at not one, but two meta-analyses regarding the effectiveness of chatbot use in language learning, with the first one by Lyu et al. published in November 2024 in the International Journal of Applied Linguistics, and one by Wang et al. published in June of the same year in the Review of Educational Research. These meta-analyses wanted to answer the following research questions:Β
- What is the overall effect of using chatbots on language learning performance?
- What is the moderating effect of other factors, such as subject, instrument, learning objectives, context, and characteristics like these?
The conduct of a meta-analysis largely follows that of a systematic review for the initial phases. Based on the research questions, the researchers would define a time period in which literature would be searched, and develop a search string to yield articles that might align with the research question. These search results would be filtered by various criteria, where papers failing to meet these criteria being excluded from subsequent analyses. Quality of studies would also be assessed, such as risk of bias, as well as various metrics of external validity. As such, the pool of included studies in meta-analyses and systematic reviews would be substantially smaller than the search results, since many of these would be irrelevant to the research question of interest.
Next, comes the synthesis part. Unlike a systematic review, a meta-analysis gets a little bit more statistically intensive. From the included studies, researchers would extract the effect sizes and use them for meta-regression, which gives us the overall effect size and direction given the findings of these primary sources. Other information may also be gathered as well, such as publication bias, which is visualized using a funnel plot and tested using stuff like Begg’s rank test and Egger’s test, and heterogeneity, which indicates the between-study similarity. One of these statistics includes the I-squared value, where the higher it is, the less agreement there is between studies, and hence a higher heterogeneity. Heterogeneity may also be tested using the Q test, though it does not really quantify the degree of heterogeneity like the I-squared value does. This would also indicate which sort of model to use, as high heterogeneity would generally require random effects models to be used in meta-regression. Subgroup analyses and sensitivity analyses may also be conducted after the main analyses to look at certain groups in the study, such as particular education levels, or particular target languages.
So, what did the meta-analyses find?
Lyu et al. yielded 31 studies with 41 effect sizes with considerable heterogeneity (I-squared value: 0.8075). Within these effect sizes, only 3 produced negative effects, 4 produced effects close to zero, while 34 produced positive effects, yielding the composite effect size of 0.608, and the weighted counterpart being 0.645. This indicated that chatbots had a significantly medium positive effect on learning performance. Further tests revealed that this effect is independent of learning outcomes as well. Furthermore, the results of this meta-analysis were not affected by publication bias. Most of the studies were on foreign language learning and in adult learners, and so the results were generally representative of this context and this study demographic.
Wang et al., however, yielded 28 studies with 70 effect sizes. These studies demonstrated a high degree of heterogeneity, like the meta-analysis conducted by Lyu et al., with an I-squared value of 0.90. The study reported an effect size of 0.484, suggesting that chatbots were effective in improving language learning performance. Egger’s test revealed that there was evidence of asymmetry, suggesting some publication bias. Further analysis showed that there was a tendency of results to report non-statistically significant findings than statistically significant findings. However, given that there were just 28 studies included in this meta-analysis, the researchers thought that this finding was ‘inconclusive’.
There were moderating effects by certain characteristics, with the meta-analysis done by Lyu et al. showing that the effects were moderated by accessibility to the chatbot, modality of the chatbot, involvement of GenAI, and nature of comparison groups in the study designs. However, meta-analysis by Wang et al. reported that educational level, language level, interface design, and interaction capability affected the overall effectiveness. A limitation to how much we can interpret from these results was the small sample size of studies, constraining the feasibility and the capability to draw meaningful interpretations for these more in-depth analyses.
It is quite interesting that both studies reported an overall positive effect of chatbot use on language learning performance relative to non-chatbot use, with moderate effect. Does that mean that you should use chatbots in your language learning regimen?
Well, not necessarily.
While the reported effect sizes are substantial, we should be cautious that these are drawn from a small number of studies and effect sizes. Furthermore, the effect sizes estimated in individual studies are also based off small sample sizes as well. As such, more synthetic studies would be needed down the line to verify these results. These results were also mainly drawn from adult learners of foreign languages, which could yield different findings from synthetic studies of other learning populations, language learning contexts, and language ability. LLMs are generally trained on existing corpora of data, and so, if oneβs target language has less available data for an LLM to train on, chances are, the effectiveness of a chatbot in this learning context would be considerably lower than a target language with an abundance of available training data, like Spanish and English. As these systematic reviews generally look at more commonly learned foreign languages, caution must be exercised if you do intend to use chatbots to learn lesser known foreign languages and some indigenous languages, where users are less willing to allow language models to train on their languages.
Another important caveat is that both systematic reviews did not quite assess chatbots as moderators of language learning effectiveness, as how they are used or integrated into the learning process may differ or overlap from learner to learner, or study to study. As such, your mileage may also vary should you choose to integrate chatbots into your language learning regimen.
Nevertheless, there are several perspectives conveyed in the reviews that I think are worth considering. One generally would learn a language with the intention of communicating with people in their target language. However, people are naturally judgmental. As such, a major obstacle a learner would face is speaking anxiety, as they might develop uneasiness of how they would be corrected by the person they are conversing with. This of course, when taken negatively, could impact the speaking confidence of language learners, and at worst, discourage them from using that language. Chatbots, however, remove this judgmental part outright, providing what users perceive as a supportive and accommodating environment to speak and use their target language beyond their lessons, generally without fear of being corrected too harshly. Then again, there will come a point where the user would have to shed these training wheels and face communicating in the real world, where there is a tendency to be judged or criticised for some errors in speaking. How then, might the users cope with this change in spoken environment? This aspect is not quite explored in either review, nor in past empirical literature, but has been a criticism against the reliance of chatbots to provide a bubble of comfort for communicating in a target language.
The use of a chatbot as a language tutor, as a result, has been one of the sticking points behind the positive attitudes demonstrated in both meta-analyses. With flexibility, ability to provide a personalised experience, giving the perception of a supportive environment to communicate, and the provision of instant feedback, these chatbots would be perceived to be worthy contenders in the grand scheme of language learning applications. However, some of the effectiveness seen in language learning here can be attributed to how chatbots are integrated in the language learning experience, rather than just the chatbot itself.
Having received these pieces of evidence, I have gathered some perspectives on why people would like to use these chatbots to learn foreign languages. However, I am quite sceptical of the effect sizes reported in both meta-analyses. As our understanding of the effects of GenAI in education and pedagogy develops over time, we might start to gather more evidence in other learner groups and learner contexts, and perhaps the picture might not be as rosy as it initially seems, like right now. Nevertheless, I will share my reflections and review of using chatbots in language learning, and give my own personal opinions if I would recommend doing so, or not.
The Meta-Analyses
Lyu, B., Lai, C., Guo, J. (2024) Effectiveness of chatbots in improving language learning: A meta-analysis of comparative studies, International Journal of Applied Linguistics, 35(2), pp. 834-851. https://doi.org/10.1111/ijal.12668.
Wang, F., Cheung, A.C.K., Neitzel, A.J., Chai, C.S. (2024) Does chatting with chatbots improve language learning performance? A meta-analysis of chatbot-assisted language learning, Review of Educational Research, 95(4), pp. 623-660. https://doi.org/10.3102/00346543241255621.