As someone fluent in English and Chinese, Wesley W. Koo has used AI-fueled large language models in both languages in his work as an assistant professor in Management and Organization at Johns Hopkins Carey Business School.
“What I found was a difference in the amount of content and the usefulness of the content that was being spit out,” he says. “For example, if I asked the LLM to summarize a paper, the English version came back with 10 to 12 points highlighted, while the Chinese version offered seven to eight points.”
The experience inspired him to embark on a study, published recently in Nature’s Scientific Reports, which examined the quality of AI-generated responses across three languages: English, Chinese, and Arabic.
“Recent research has suggested that multilingual AI technologies can elevate the skills of low-skilled workers. But what has been understudied is a human-centric perspective focused on the effectiveness of AI-generated content in human work and the linguistic context in which the work takes place,” Koo says.
In his two-part study, Koo found that when human participants used AI-generated content in email writing, it was associated with “less actionable and less creative” work output in Arabic and Chinese than in English. Moreover, non-English participants were particularly disadvantaged when dealing with more technical tasks, such as product design or scientific discovery.
“The most popular LLMs — such as OpenAI’s ChatGPT, Google’s Gemini, and Anthropic’s Claude — were trained primarily on English data, and most non-English speakers tend to use such English-based products. Even the Chinese DeepSeek model is heavily trained on English content,” notes Koo. “The findings of my study underscore the need for more language-inclusive LLMs to support workers worldwide, since current technologies run the risk of widening the productivity gap between English and non-English speakers, particularly in more technical domains.”
Less actionable, less creative
In the first part of his study, aimed at evaluating content, Koo focused on the quality of AI-generated responses to common business questions in all three languages. He used ChatGPT-4 to supply 200 prompts to common business-related issues, which were translated into English, Chinese, and Arabic, resulting in 600 AI-generated responses. He enlisted a team of six native speakers, two for each language, to evaluate the responses, following a rigorous two-week training program aimed at providing consistency in the evaluation criteria.
Evaluation showed that compared to English responses, responses in Arabic were notably lower in the four areas under study: completeness, relevance, actionability, and creativity. Chinese responses were slightly better but still lower than English responses in all four areas.
”While AI is touted to have the potential to elevate workers in less advantaged positions, this research shows that it could reinforce the dominant position of advantaged populations."
A wider gap for technical tasks
In the second part of the study, which Koo describes as the “core” of his project, evaluators examined the quality of the emails that participants wrote in response to AI output, as well as their productivity and perceptions.
Here, Koo found that the use of AI led to longer emails written in English than in Arabic or Chinese.
In gauging participants’ self-reported confidence about the quality of their output and perceived difficulty of the email-writing task, he found that Chinese individuals were more likely to embrace and feel confident about AI. He says this aligns with recent studies showing that people in Eastern countries are more optimistic about the usage of AI. But he cautions that “more confidence with AI does not necessarily lead to better work with AI-generated content”