Introduction

With the exponential development of information technologies, artificial intelligence (AI) technologies have been widely applied to various fields, including education (Airaj, 2024). Recently, GenAI, which integrates natural language processing and visual content generation into educational processes, has attracted the keen attention of researchers and practitioners in education due to its powerful data generation abilities, content creation capacity, and model recognition capacity. GenAI offers potential opportunities for education, which may be greatly impacted and revolutionized (Longhurst et al., 2020). In the field of education, GenAI functions in various forms (Li et al., 2024). It may offer accurate teaching feedback for teachers to adjust their teaching strategies and optimize their teaching effectiveness by intelligently analyzing students’ learning behaviors and data on their academic achievements (Richmond and Nicholls, 2024).

Research gap in the literature

The main research gap in the literature is that few studies have been committed to exploring the pooled effect of GenAI technologies on educational outcomes, such as academic achievements, higher-order thinking abilities, writing skills, and few studies have explored game-assisted GenAI and differences in the effect of GenAI across countries and educational levels (see Table 1 for details). Despite the promising application of GenAI technologies in education, there are still controversial findings. Some studies have reported that GenAI may significantly improve students’ learning engagement, self-regulation, motivation, and autonomous learning (Han et al., 2025), while others argue that over-reliance on GenAI technologies may negatively impact students’ critical cognitive capabilities, e.g., decision-making, critical thinking, and analytical reasoning (Zhai et al., 2024). Therefore, it is necessary and meaningful to objectively evaluate the effect of GenAI on educational outcomes.

Table 1 A comparative analysis between the published research and the current study.

To bridge the research gaps, the current study will provide a comprehensive understanding of the effect of GenAI on educational outcomes, the added value of GenAI feedback and game-assisted GenAI, and the moderating effects of countries and educational levels.

As shown in Table 1, prior meta-analyses and systematic reviews have investigated the effect of GenAI on educational outcomes to a relatively narrow extent. They pivoted on ChatGPT’s role in writing instruction and feedback (Ibrahim & Kirkpatrick, 2024), addressed policy, limitations, and proactive measures (Ogunleye et al., 2024; Salman et al., 2025), or focused on ethical and legal issues (Al-Shabandar et al., 2024). However, this study departs from these narrow inquiries by pooling GenAI’s effects on essential educational outcomes, such as academic achievements, higher-order thinking, and writing skills. It will also examine under-explored elements, such as GenAI feedback, game-assisted GenAI, and the moderating effects of countries and educational levels.

Research questions

To bridge the above-mentioned research gaps, this study aims to address the following research questions:

RQ1. May GenAI technologies outperform non-GenAI technologies in terms of academic achievements, higher-order thinking abilities, and writing skills?

RQ2. May GenAI feedback outperform non-GenAI feedback in terms of educational outcomes?

RQ3. May game-assisted GenAI outperform non-game-assisted GenAI in terms of educational outcomes?

RQ4. Are there significant differences in the effects of GenAI technologies on educational outcomes across different countries (China, Pakistan, Korea, Turkey)?

RQ5. May GenAI technologies exert significant positive effects on educational outcomes at both university and secondary levels?

This study attempts to systematically review and synthetically analyze the effect of GenAI/feedback on educational outcomes, including academic achievements, higher-order thinking abilities, and writing skills, the effect of game-assisted GenAI on educational outcomes, and differences in the effect of GenAI on educational outcomes across countries and educational levels. Meta-analysis is a statistical approach that quantitatively pools the effect sizes of independent studies to enhance the generalizability and reliability of conclusions (A. Chen et al., 2025). Meta-analysis enables researchers to evaluate the effect of GenAI on educational outcomes and reveal the benefits and challenges of GenAI, thus offering constructive suggestions for researchers, practitioners, and policymakers.

Theoretical framework

GenAI may improve academic achievements, provide personalized learning experiences, facilitate teaching effect, promote teacher professional development, integrate gamification elements, and adapt to different contexts across various countries. GenAI technologies are expected to improve student engagement and learning outcomes (Francis et al., 2025) due to powerful natural language processing abilities and machine learning functions. GenAI technologies may offer personalized learning experiences (Seo et al., 2024) and effective feedback to improve learning outcomes (Corbin et al., 2025). By identifying the weaknesses of students, GenAI may improve content creation, assessment, feedback, and the design of learning activities (Moundridou et al., 2024). GenAI may also help teachers and educators adjust their teaching styles and modify educational content to meet the diverse needs of different students and promote teachers’ professional development (Aguilar-Cruz and Salas-Pilco, 2025).

GenAI technologies may process and analyze a large amount of data, identify learning models, and provide timely and contextual feedback (Dann et al., 2024). The integration of gaming factors into learning via GenAI technologies may enhance real-time student interactions (Moon et al., 2025). GenAI plays an important role in revolutionizing teaching practices, adaptive learning, and student engagement (Hirpara et al., 2025). By investigating the effect of GenAI on academic achievements and cognitive skills and exploring GenAI feedback, game integration, cross-cultural adaptability, and various educational levels, this study attempts to provide constructive guidance and recommendations for future theoretical research and teaching practices.

Literature review

Definitions of variables

Academic achievements are operationally defined in this study as the mastery of knowledge and learning effectiveness in a given field through tests, assignments, examinations, and standardized assessments, where basic conceptions, formulas, and theoretical elements are involved (Moore, 2019). Academic achievements are considered the most direct, traditional core indicator of GenAI-driven educational outcomes.

Higher-order thinking refers to the complicated cognitive abilities and those beyond basic memory (Liu et al., 2024). Higher-order thinking abilities involve three components, i.e., problem-solving, creativity, and critical thinking (Liu et al., 2024), which are reflected in the form of analysis of complicated problems, critical analysis of information, and innovative problem-solving abilities based on constructed knowledge structures.

Writing skills are operationally defined as the comprehensive abilities of transferring information and expressing viewpoints using languages, which include logical cohesiveness, expression accuracy, and content innovation (Chang et al., 2024). Logical cohesiveness focuses on whether the structure is clear and whether arguments are well-matched with evidence. Expression accuracy evaluates norms of grammar and appropriateness of lexical selection. Content innovation highlights the uniqueness of viewpoints and the creation of perspectives (Barroga and Matanguihan, 2021). Three core dimensions of writing skills may reflect the application skills of language and the logic of thinking.

Rationales for the inclusion of the above variables as educational outcomes

The core educational objective is to integrate knowledge transfer into the cultivation of abilities (Jackson et al., 2019). The effectiveness of knowledge transfer is reflected through academic achievements. The cultivation of abilities is demonstrated in the development of higher-order thinking abilities, as well as cognitive and non-cognitive skills (Peng et al., 2021). The powerful compatibility of GenAI technologies may enhance personalized learning effectiveness (Han et al., 2025). The technological dimensions, such as natural language processing and data-driven deep learning (Khan et al., 2023), interact with each other to improve academic achievements by providing rich personalized learning resources. In the research on modern educational technologies, reliable and valid instruments have been developed to measure academic achievements, higher-order thinking abilities, and writing skills (Wu et al., 2024). The dimensions have been frequently regarded as an indicator of educational evaluation across different countries and on various educational levels, thus enhancing the representativeness of cross-sample analyses in different contexts.

Core gaps in existing literature: a thematic synthesis

While international researchers have explored GenAI-driven educational outcomes, very few conduct thematic integration across studies. Most previous studies consider research results as isolated pieces rather than integrated patterns—gaps this study explicitly addresses. For one thing, Zhang et al. (2025) examined inconsistencies in the effect of GenAI technologies on writing performance (e.g., grammar use and sentence variety) and controversial issues over the impact of GenAI technologies on higher-order skills. However, no synthetic analysis connects these controversies to types of GenAI treatments, GenAI tools, GenAI feedback, or game-assisted GenAI. For another, Sun and Lan (2025) revealed that Chinese EFL learners studied the effect of GenAI on educational outcomes through sociocultural factors. However, this local study failed to connect to a broader context in the use of GenAI, which prevented stakeholders from comprehensively understanding the effect of GenAI technologies on educational outcomes. Therefore, it is necessary to carry out a unified, evidence-based assessment, which will be addressed in this study.

Another noteworthy research gap lies in the focus of existing studies on the description of issues, such as GenAI feedback, game-assisted GenAI, regional variances, and educational levels. But most of them focus on positive or negative outcomes rather than identifying the root causes of these findings. For example, it was found that GenAI feedback was strongly connected to technical characteristics such as timeliness and personalization (Li et al., 2025; Dann et al., 2024). However,the theoretical frameworks, such as meta-cognition theory (Aburayash, 2021), failed to be revealed to account for feedback-shaped learning. This omission indicates that stakeholders did not distinguish the differences between GenAI feedback and the effect of GenAI feedback on learning outcomes. Another example was that game-assisted GenAI research reported both the effect of game-assisted GenAI on educational outcomes (Chien et al., 2024) and distraction factors (Ramos and Melo, 2019), but failed to provide insights into the core reasons, such as game elements and learning personalities. Without this integrative analysis, the literature remains a collection of isolated pieces, rather than a meaningful and integrative insight.

Further synthesized analysis is necessary when regional and educational level differences are included. For instance, the use of GenAI technologies in both China and Pakistan aimed to address the inequity of educational resources (Naz et al., 2019; Ahmad et al., 2023). On the contrary, cultural trends towards teacher-centered pedagogy slowed down the GenAI progress in Korea and Turkey (So et al., 2024; Yilmaz et al., 2024), where printed curricula might not be compatible with GenAI-assisted approaches. However, few studies connect regional differences to GenAI’s educational effects, leading to the lack of effective guidance for policymakers. Furthermore, GenAI technologies were explored in the context of universities (Lan et al., 2025) and secondary schools (Ng et al., 2024), but no moderating factors, such as academic rigor and learning interest, were explored. This research gap leads to the failure in identifying the effect of GenAI technologies on different educational levels and regional differences, which will be bridged by the current study.

Academic achievement, higher-order thinking, and writing skills

GenAI enhances academic achievement, higher-order thinking, and writing skills when used as a “cognitive aid,” but negatively influences them in case of over-reliance upon GenAI as a replacement for human teachers (Han et al., 2025). This pattern clarifies the contradictory findings in the literature. Blanca Ibanez et al. (2020) reported that GenAI technologies could improve academic achievement via personalized learning resources and timely feedback, echoed by Nair and James (2022). GenAI technologies could separate difficult issues into pieces of easier problems, enhancing higher-order thinking skills, such as critical thinking abilities and analytical skills (Borge et al., 2024). The real-time feedback and multiple explanation functions could assist university students to better understand professional knowledge and improve their academic achievements (Wang et al., 2024). Over-reliance on GenAI is highly frequent among middle school students (Ng et al., 2024). This divergence underscores the importance of human instruction rather than the replacement of GenAI technologies in educational practices.

GenAI feedback

GenAI-assisted instruction may improve timeliness and personalization of feedback compared with non-GenAI instruction, though student distrust and emotional factors reduce its effectiveness in some contexts. GenAI-assisted feedback could adjust the contents of feedback, catering to learners’ needs (Xu et al., 2025; Dann et al., 2024). Xu et al. (2025) revealed that GenAI feedback could provide accurate suggestions for students to improve their learning outcomes, compared with teacher feedback. Dann et al. (2024) also echoed that GenAI could provide timely and rapid feedback to avoid delays. Human teachers did not have enough time to create detailed and personalized feedback (e.g., Lin and Crosthwaite, 2024). However, Cosentino et al. (2025) revealed no significant differences between GenAI and non-GenAI feedback in terms of cognitive load or visual processing, while Henderson et al. (2025) argued that students trusted GenAI feedback less than human teacher feedback.

Game-assisted GenAI

Game-assisted GenAI’s impact on educational outcomes heavily relies on two independent factors, i.e., game design in line with educational goals and learner age engaged in game-assisted education. Whether or not the game is properly designed exerts a great influence on the role of AI in educational outcomes. Game-integrated GenAI could enhance learning engagement and educational outcomes if games were designed to enhance learning or provided rewards for completion of an assignment (Chien et al., 2024). It was also reported that game-assisted GenAI could improve educational outcomes, especially among older learners, because gaming elements could trigger learner motivation and encourage timely task completion (Moon et al., 2025). However, entertainment-oriented games could distract learners’ attention and reduce knowledge retention (Ramos and Melo, 2019). Game-integrated GenAI has been under-explored, and few studies linked this under-exploration to game design quality (Chen et al., 2022).

GenAI in regional contexts (China, Pakistan, Korea, Turkey)

Educational resource equity and cultural attitudes toward pedagogy primarily drive regional variations in GenAI’s effectiveness, pooling the fragmented regional differences in GenAI’s effectiveness. The two factors may also explain regional differences in the effect of GenAI on educational outcomes, integrating fragmented findings across countries.

Regional differences in the effect of GenAI on educational outcomes are closely related to educational needs and cultural factors. GenAI could consistently outperform non-GenAI technologies in education because the former could effectively address the inequity in the distribution of educational resources in China and Pakistan (Naz et al., 2019; Ahmad et al., 2023), where high-quality educational resources could be distributed to remote areas. Besides, GenAI should also be designed in accordance with cultural priorities, e.g., cultural emphasis on knowledge accumulation (Casillo et al., 2022). However, the effect of GenAI technologies on educational outcomes proved negative in Korea and Turkey, possibly due to their cultural preferences for traditional teacher-student interaction (So et al., 2024), curricular misalignment (Yilmaz et al., 2024), and technical limitations, such as algorithm bias and low-quality training data (Karimova et al., 2025). So et al. (2024) believed that the interactive modes of GenAI technologies could weaken teachers’ authority, which is a core value in Korean universities. GenAI-based learning material tended to deviate from the guidelines of national curricula, leading to confusion and ineffectiveness in academic performance (Yilmaz et al., 2024). These disconnections accounted for the negative effect of GenAI technologies on educational outcomes.

GenAI in universities and secondary schools

GenAI improves outcomes at both university and secondary school levels, but its application strategies vary due to learner requirements and learning goals, which is an essential key distinction absent from previous research.

Specific learning goals should be considered when GenAI-assisted pedagogy is designed for different levels of education. GenAI could enhance research skills and academic writing by providing personalized learning resources (Lan et al., 2025; Wang et al., 2024), such as literature summaries and data analysis tools, catering to university students’ needs for independent and rigorous inquiry at the university level. Lan et al. (2025) revealed that university students could speed up their literature review work by using GenAI technologies and cited more papers than those without the assistance of GenAI technologies. It was also reported that the data visualization tools in GenAI technologies could improve the analysis quality of academic papers (Wang et al., 2024). At the secondary level, GenAI could enhance learning engagement and basic knowledge acquisition through interactive AI tools, such as adaptive quizzes and visual learning presentations (Ng et al., 2024; Delaney et al., 2025). It is noteworthy that university students tended to possess higher AI literacy than secondary students, which made them more effective in using GenAI technologies in education.

Research hypotheses

This study proposes the following alternative hypotheses on the basis of the thematic analysis of existing literature, as well as the research gaps identified:

H1. GenAI technologies may cause higher academic achievements, higher-order thinking abilities, and writing skills than traditional technologies.

H2. GenAI feedback may cause significantly higher educational outcomes than non-GenAI feedback.

H3. Game-assisted GenAI may cause significantly higher educational outcomes than non-game-assisted GenAI.

H4. GenAI technologies may cause significantly higher educational outcomes than non-GenAI technologies across countries (China, Pakistan, Korea, and Turkey), though the magnitude of these effects may vary by regional context.

H5. GenAI technologies may cause significantly higher educational outcomes than non-GenAI technologies at both university and secondary levels.

Transition to methods

This study quantitatively tests the research hypotheses and bridges the research gaps based on the theoretical framework, literature synthesis, and research hypotheses proposed above. Below are the detailed research methods, ranging from literature search, screening, data retrieval, to analytical approaches, as well as the methods to ensure the rigor and reliability of the findings.

Research methods

This study aims to comprehensively explore the effect of GenAI on educational outcomes and test the proposed hypotheses through meta-analysis and systematic review methods based on the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) protocol (Farrus, 2023).PRISMA, aiming to improve the transparency and reliability of research, is a framework frequently used to conduct a systematic review and meta-analysis. The PRISMA protocol consists of 27 checklist items and a four-phase flow chart, covering the proposal of research problems, literature search and filtering, data retrieval, quality evaluation, and result reporting. Adherence to PRISMA may enhance the reliability and validity of the research, resulting in convincing results.

Literature search

Based on the framework of PICO (Population, Intervention, Comparison, Outcome) (Wu et al., 2024), we clarified the research subjects (students, teachers, or educational institutes in educational contexts), and compared GenAI technologies with non-GenAI technologies, GenAI feedback with non-GenAI feedback, game-assisted GenAI and non-game-assisted GenAI.

Selection of databases

To obtain related studies across various countries and areas, we searched comprehensive online databases, including Scopus, Web of Science, Springer, Wiley, and ScienceDirect, covering the period from their inception to 2025. We also searched gray literature, such as dissertations and conference proceedings, to minimize publication bias.

In March 2025, we constructed multi-faceted search terms, including GenAI (e.g., generative AI, generative artificial intelligence, large language models, and ChatGPT), educational outcomes (e.g., academic achievement, higher-order thinking abilities, and writing skills), feedback (e.g., GenAI feedback, non-GenAI feedback), assistance methods (e.g., game-assisted GenAI and non-game-assisted GenAI, country (e.g., China, Pakistan, Korea, and Turkey), and educational levels (e.g., university, secondary school, and primary school). We flexibly arranged the search terms using Boolean logical operators, such as “AND” and “OR”, e.g., (generative AI OR generative artificial intelligence) AND (educational outcomes OR academic achievement) AND (China OR Pakistan).

On March 9, 2025, we obtained 568 results from All Databases of Web of Science for: “artificial intelligence” OR “large language model*“ OR ChatGPT OR “generat* AI”(Topic)AND educat*(Topic)AND “control group*“(Topic). We obtained 472 results from Scopus for: “artificial intelligence” OR “large language model*“ OR ChatGPT OR “generat* AI“AND educat*AND “control group*“in title, abstract, or keywords. By searching Springer Nature Link, we obtained 15 results for: “artificial intelligence” OR “large language model*“ OR ChatGPT OR “generat* AI“AND educat*AND “control group*“ in keywords. We obtained 24 results from ScienceDirect for: educationAND “control group” in Title, abstract, keywords, and “artificial intelligence” OR “large language model” OR ChatGPT OR “generative AI” in Title. We retrieved 17 results from Wiley for “artificial intelligence” OR “large language model” OR ChatGPT OR “generative AI” in Title and “education “AND “control group” in Abstract. Finally, this initial search yielded a total of 1096 records across all databases.

Literature screening

Two independent raters screened the literature based on the protocol of PRISMA. They first included and excluded the literature based on the inclusion and exclusion criteria. If they could not reach an agreement on any selection, a third rater would be invited to make a final decision (See Fig. 1 for details). The inclusion criteria are: (1) The researchers are students, teachers, or educational institutes who use GenAI in educational contexts; (2) The literature reports the quantitative or qualitative data related to the research hypotheses proposed in this study, e.g., academic achievements, higher-order thinking skills, or writing skills; (3) The research designs belong to experimental, quasi-experimental, or survey studies, e.g., game-assisted GenAI vs. non-game-assisted GenAI, GenAI feedback vs. non-GenAI feedback, and GenAI-based educational outcomes in different countries. The exclusion criteria are: (1) The themes are not related to this study; (2) unreliable research designs: no control group, no reliable data, or no reliable intervention; (3) key data, e.g., mean, standard deviation, or sample size, are not revealed; (4) unpublished documents, retracted or overlapped publications; (5) The documents are news items, patents, editorial material, awarded grant, or opinion studies. We finally included 53 results for the meta-analysis (See Table 2a–d for details).

Fig. 1
Fig. 1
Full size image

A Literature filtering flowchart based on the PRISMA protocol.

Table 2 Details of included studies.

In the included studies, the instruments used to measure educational outcomes are homogeneous to some extent. Academic achievements are measured mainly through standardized examinations, adapted tests, and assignments. Higher-order thinking skills are measured by structured evaluation tools, e.g., question answer rubrics, questionnaires, and innovative problem-solving criteria. Writing skills are measured by writing quality evaluation criteria, including logic cohesiveness, accuracy, and content innovation. The valid and reliable instruments are preferred to ensure data reliability and rigor.

The screening process involved four specific steps, including initial deduplication, automation tool pre-screening, abstract screening, and full-text eligibility assessment. Initial deduplication process removed 9 duplicate records from a total of 1096 results, leading to 1087 records. The automation tool pre-screening process marked 13 records as ineligible, reducing the pool to 1074 records. The abstract screening was carried out by two independent raters who reviewed the abstracts of the 1074 records, excluding 391 documents beyond the scope, leaving 683 records for full-text assessment. The full-text eligibility assessment process screened the 683 full-text records based on the inclusion/exclusion criteria. The 630 records were excluded due to unreliable research designs (n = 205), lack of key data (n = 267), unpublished/retracted/overlapped publications (n = 31), and news items, patents, editorials, grants, or opinion studies (n = 127). We finally included 53 studies for the current study.

Summary of included studies by category

To facilitate reading, the included studies were categorized as GenAI Feedback, Game-Assisted GenAI, Country-China & Pakistan, Country-Korea & Turkey, University-Level Education, Secondary-Level Education, Academic Achievements Focus, Higher-Order Thinking Focus, and Writing Skills Focus (Table 3).

Table 3 Summary of included studies by thematic category.

The quality of the included literature is evaluated based on Standards for Reporting on Empirical Social Science Research in AERA Publications: American Educational Research Association (2006). The evaluation standards include Problem Formulation, Design and Logic, Sources of Evidence, Measurement and Classification, Analysis and Interpretation, and Ethics in Reporting. Each standard accounts for 1 to 3 points, totaling 6 to 18 points. Studies were specifically categorized into the following ranks based on the quality scores:

High-quality studies (12–18 points): the studies with high quality feature systematic problem formulation, rigorous design and logic, solid sources of evidence, reasonable measurement and classification, unbiased analysis and interpretation, and ethical reporting (n = 28, 52.8% in the included studies).

Medium-quality studies (9–11 points): the studies with medium quality satisfy basic academic research requirements but have minor shortcomings, such as small sample sizes and limited data availability (n = 19, 35.8% in the included studies).

Low-quality studies (3–8 points): the studies with low quality fail to meet basic academic standards due to a lack of rigorous design or incomplete analysis/interpretation) (n = 6, 11.4% in the included studies).

A pilot test was carried out on 10 randomly selected studies to calculate their scores before formal evaluation, which achieved a satisfactory inter-rater reliability coefficient (Cohen’s kappa = 0.82). If both raters could reach an agreement even after discussion, a third rater would be invited to make a final decision. The raters included or excluded the studies based on various factors, such as research design adequacy, sampling quality, instrument reliability, publication bias risks, and ethical compliance. These factors were included to cross-check the quality of the literature to ensure comprehensive quality appraisal.

Data retrieval

We retrieved the data related to the research objectives. The contents of the retrieved data included author, publication year, mean, standard deviation, sample size, educational level, research geography, educational outcome, feedback type, assistant model, and country or area where the study was conducted.

Literature quality evaluation

Two independent raters first evaluated the quality of the research designs. They evaluated if the research design could address all the research questions, if the control group was properly established, and if blinding methods were properly implemented. They then evaluated the sampling quality. The sample should be able to represent the population, which requires scientific sampling techniques, such as random and stratified sampling. The sample size should be large enough to ensure the results are robust. The instruments to collect data should be reliable. The collection process should be normalized to obtain complete and accurate data. The research results should be reliable and stable, leading to reasonable and logical conclusions. Other evaluation metrics include publication bias evaluation, research fund, interest conflicts, and ethics issues.

Interventions

The interventions were classified as use of GenAI technologies, GenAI feedback, and game-assisted GenAI. GenAI technologies may autonomously generate text, image, and sound, which may be used in education to design teaching materials, create personalized exercises, and customize learning data, catering to different teaching and learning needs. This sort of intervention is carried out based on the constructivist theory, positing that learning is a process where learners actively construct knowledge rather than the passive reception of knowledge (Brandon and All, 2010).

GenAI feedback is operationally defined as the automatically generated comments and suggestions via GenAI tools for tasks, such as assignment, question answering, and oral expression. This intervention works based on the theory of meta-cognition (Aburayash, 2021), strengthening students’ abilities of monitoring, evaluating, and regulating the learning process. The GenAI feedback may not only correct errors, but also guide students’ reflection process, helping students to construct their own cognitive levels and self-regulate their learning strategies.

The game-assisted GenAI intervention integrates games into GenAI technologies. GenAI may generate English word games and adventuring contexts, where GenAI may modulate the difficulties of game-assisted learning, generate individualized feedback, and combine entertainment with personalized learning. This intervention is developed on the basis of Self-Determination Theory (Chang, 2026), which argues that the inner motivation will be greatly improved if individual self-direction, competence perception, and belonging sense are satisfied.

Data analyses

Calculation of effect sizes

We first calculated the mean differences, and adjusted the initial effect size to obtain the small-sample effect size (Hedge’s g) (Carnevali et al., 2024) using the formula: 1 - 3/(4×total sample size). Effect size thresholds: small (g ≈ 0.2), moderate (g ≈ 0.5), and large (g ≈ 0.8) (McAloon and Armstrong, 2024).

Tests of heterogeneity

Cochrane Q tests and I² statistics were adopted to measure the degree of heterogeneity (Dai et al., 2024). If I² < 50%, the heterogeneity will be considered statistically insignificant, and the fixed-effect model will be adopted for the meta-analysis. If I² > 50%, the heterogeneity will be considered statistically significant, and we will adopt a random-effect model for the meta-analysis (A. Chen et al., 2025).

Subgroup analysis

We conducted the subgroup analyses in terms of academic achievements, higher-order thinking abilities, writing skills, GenAI feedback, game-assisted GenAI, countries, and educational levels.

Sensitivity analysis

We conducted the sensitivity analysis via a leave-one-out meta-analysis (Dondio et al., 2023). We conducted the meta-analysis by removing a study one by one and then observed the pooled effect sizes, as well as their stability and robustness. If, after one study was removed, the effect size significantly changed, the results would not be considered stable and robust.

Publication bias

We tested the publication bias through funnel plots, Egger’s test, and Begg’s test (Lozano-Blasco and Cortes-Pascual, 2020). If there is the presence of publication bias, trim-and-fill analysis will be used to modify it (Wu et al., 2023).

Qualitative analysis

We conducted the qualitative analysis based on the included qualitative studies using the thematic analysis method. Firstly, we coded the text data and classified the same contents into a category for further analysis. We then complemented the quantitative analysis by retrieving and analyzing the themes, such as the benefits, challenges, and educational outcomes in the application of GenAI in different countries.

Results

Heterogeneity testing

We drew a Galbraith plot to test whether there is heterogeneity in the effect sizes of included studies. As shown in Fig. 2, the x-axis indicates the precision of standard errors of estimated effect sizes, while the y-axis indicates the standardized effect sizes (Hedge’s g). The red line is the regression line, which crosses the zero point, indicating the fixed-effect model. The shade alongside the regression line indicates the 95% confidence interval, while a dot indicates an effect size. No heterogeneity is revealed if all the dots are located within the confidence interval. However, heterogeneity will be present if there are any dots outside the confidence interval. It is clearly shown that numerous dots fall outside the confidence interval, indicating the presence of heterogeneity. Therefore, we adopted a random-effect model to conduct the meta-analysis.

Fig. 2
Fig. 2
Full size image

A Galbraith plot to test the heterogeneity.

H1. GenAI technologies may cause higher academic achievements, higher-order thinking abilities, and writing skills than traditional technologies

As shown in Fig. 3, we conducted a meta-analysis via a random-effect model in terms of academic achievements, higher-order thinking abilities, and writing skills in GenAI-assisted education. The obtained effect sizes are significantly heterogeneous (I² > 50%, p < 0.05). We thus adopted a random-effect model to conduct the meta-analysis. GenAI-assisted education may cause significantly higher academic achievements (g = 0.40, small-to-moderate effect; 95%CI = 0.08 ~ 0.71), higher-order thinking abilities (g = 0.72; 95%CI = 0.30 ~ 1.13, moderate-to-large effect), learning motivation (g = 0.81; 95%CI = 0.20 ~ 1.41, large effect), and writing skills (g = 0.76; 95%CI = 0.18 ~ 1.35, moderate-to-large effect) than the non-GenAI-assisted educational approach. Therefore, we accept the alternative hypothesis that GenAI technologies may cause higher academic achievements, higher-order thinking abilities, and writing skills than traditional technologies.

Fig. 3
Fig. 3
Full size image

Meta-analysis results in academic achievements, higher-order thinking abilities, learning motivation, and writing skills.

H2. GenAI feedback may cause significantly higher educational outcomes than non-GenAI feedback; H3. Game-assisted GenAI may cause significantly higher educational outcomes than non-game-assisted GenAI

Figure 4 clearly shows the effect of GenAI feedback and game-assisted GenAI on educational outcomes through a forest plot. In the forest plot, sample sizes, means, and standard deviations are shown, together with effect sizes and 95% confidence intervals. The indicators of heterogeneity, including I², z, and p values, are also shown in the plot. We adopted a random-effect model for the meta-analysis due to the presence of heterogeneity of effect sizes (I² > 50%, p < 0.05). It is demonstrated that GenAI feedback may exert a significant influence on educational outcomes (g = 1.27; 95%CI = 0.49 ~ 2.05), while game-assisted GenAI may not cause significant differences in educational outcomes (g = 0.24; 95%CI = -0.40 ~ 0.88). Therefore, we accept the second hypothesis: GenAI feedback may cause significantly higher educational outcomes than non-GenAI feedback, while rejecting the third hypothesis: Game-assisted GenAI may cause significantly higher educational outcomes than non-game-assisted GenAI.

Fig. 4
Fig. 4
Full size image

Meta-analyses of GenAI feedback and game-assisted GenAI.

H4. GenAI technologies may cause significantly higher educational outcomes than non-GenAI technologies across countries (China, Pakistan, Korea, and Turkey), though the magnitude of these effects may vary by regional contexts

To measure the effects of GenAI technologies on educational outcomes across different countries, we designed and conducted a meta-analysis. The subgroup includes different countries, including China, Korea, Turkey, and Pakistan. We adopted a random-effect model for the meta-analysis due to the presence of significant heterogeneity in most of the obtained effect sizes (I² > 50%, p < 0.05). As shown in Fig. 5, in China (g = 0.71; 95%CI = 0.5 ~ 0.92) and Pakistan (g = 0.75; 95%CI = 0.58 ~ 0.92), GenAI technologies cause significantly higher educational outcomes than the non-GenAI technologies. However, in Korea (g = 1.68; 95%CI = -0.41 ~ 3.78) and Turkey (g = 0.01; 95%CI = -0.23 ~ 0.25), no significant differences have been revealed in the effect of GenAI technologies on educational outcomes. Therefore, we accept the hypothesis that GenAI technologies may cause significantly higher educational outcomes than non-GenAI technologies in China and Pakistan, while rejecting the hypothesis that GenAI technologies cause significantly higher educational outcomes than non-GenAI technologies in Korea and Turkey.

Fig. 5
Fig. 5
Full size image

Meta-analysis results in different countries.

H5. GenAI technologies may cause significantly higher educational outcomes than non-GenAI technologies at both university and secondary levels

To test the effects of GenAI technologies on educational outcomes at different educational levels, we adopted a random-effect model considering the presence of heterogeneity in effect sizes (I2 > 50%, p < 0.05). At both university (g = 0.70; 95%CI = 0.52 ~ 0.88) and secondary (g = 0.80; 95%CI = 0.15 ~ 1.44) educational levels, GenAI technologies cause significant differences in educational outcomes (Fig. 6). Therefore, we accept the alternative hypothesis that GenAI technologies may cause significantly higher educational outcomes than non-GenAI technologies at both university and secondary levels.

Fig. 6
Fig. 6
Full size image

Meta-analysis results at different educational levels.

In general, GenAI technologies may significantly improve academic achievements, higher-order thinking abilities, and writing skills. However, GenAI technologies may not significantly improve educational outcomes in Korea and Turkey, which is contrary to the results in China and Pakistan. Compared with non-AI generative feedback, AI generative feedback may significantly improve educational outcomes. Game or gamification does not play an important role in GenAI-assisted education since game-assisted GenAI may not cause significantly higher educational outcomes than non-game-assisted GenAI. Educational levels do not cause significant differences in GenAI-assisted educational outcomes since GenAI technologies may cause significantly higher educational outcomes than non-GenAI technologies at both university and secondary levels. We summarize the research findings in Table 4.

Table 4 A summary of research findings.

To test the publication bias of the included studies, we adopted Egger’s test, Begg’s test, and the non-parametric trim-and-fill analysis. The regression-based Egger test for small-study effects indicates the presence of publication bias (z = 8.53, p < 0.01). Similarly, Begg’s test for small-study effects finds the presence of publication bias (z = 4.41, p < 0.01). The non-parametric trim-and-fill analysis of publication bias reveals a total of 116 studies, where 92 studies (blue dots) are observed, and 24 are imputed (yellow dots) (Fig. 7). Both the observed studies (g = 0.758, 95%CI = 0.556 ~ 0.961) and observed/imputed studies (g = 1.091, 95%CI = 0.885 ~ 1.298) indicate the significant differences between the GenAI-assisted group and the control group. The imputed results are consistent with the observed results, indicating the robustness and stability of the meta-analysis results.

Fig. 7
Fig. 7
Full size image

The non-parametric trim-and-fill analysis results.

A sensitivity analysis was carried out through a leave-one-out meta-analysis, which was used to examine the reliability and stability of the pooled results. The leave-one-out meta-analysis calculates and pools the remaining results after one study is removed from the original dataset. The process repeats until all the studies have been removed one by one for the pooled meta-analysis. As shown in Fig. 8, a dot indicates the effect size of a study, while the horizontal line indicates the confidence interval, followed by effect sizes (g) and specific lower and upper bounds of confidence intervals. If all the dots fall within the scope of pooled confidence intervals, the pooled meta-analysis results will be considered reliable and stable. On the contrary, any dot beyond the pooled confidence interval will indicate instability and unreliability. Figure 8 shows a reliable and stable meta-analysis result since all the effect sizes fall within the scope of the pooled confidence intervals.

Fig. 8
Fig. 8
Full size image

Leave-one-out meta-analysis results.

Discussion

Academic achievements, higher-order thinking abilities, and writing skills

GenAI may significantly improve academic achievements, higher-order thinking abilities, and writing skills. GenAI possesses the powerful ability to assemble information and provide learning resources (Almuhanna, 2024), from which students may obtain rich knowledge to solve problems in their learning activities. They may thus understand the knowledge in a deep manner and perform well in their academic pursuits. GenAI may guide students to develop in-depth thoughts and analyses. For instance, GenAI may offer multifaceted perceptions and analysis frameworks to understand complicated problems, arouse students’ thinking abilities, and enhance their higher-order thinking abilities (Borge et al., 2024), which is in line with the constructivism theory (Brandon and All, 2010) focusing on active knowledge construction rather than passive reception. GenAI may also improve students’ writing skills (Nair and James, 2022) in many aspects, such as grammar correction, vocabulary recommendation, and structure optimization, to enable formal, fluent, and logical writing, which aligns with the input hypothesis (Krashen, 1982), arguing that comprehensive input may help learners to enhance their writing skills rather than overwhelm them. This is consistent with our finding that GenAI technologies may enhance writing skills. GenAI may also provide excellent writing samples and concepts to stimulate students’ writing interest and improve their writing skills.

However, inconsistent findings have also been revealed. Over-reliance on GenAI technologies could exert negative influences– on higher-order thinking skills and critical thinking abilities, possibly harming students’ mental health and social skills. The inconsistent findings may have been caused by different applications of GenAI. In the current study, GenAI technologies, as assistant tools, provide students with personalized learning resources, multiple analysis frameworks, and teaching guidance. However, GenAI technologies have not replaced students’ self-directed thinking process in the opinion of Roe and Perkins (2024), who argued that GenAI played a role as a “cognitive aid” rather than a replacement to protect or enhance learners’ thinking abilities, while Zhai et al. (2024) revealed negative results due to over-reliance on GenAI technologies. The studies yielding negative results regarding critical thinking abilities may be caused due to direct educational outcomes through over-reliance on GenAI technologies without self-directed insights and autonomous thinking. Besides, the heterogeneous demographic information of participants, such as age, educational backgrounds, and types of GenAI tools may have contributed to conflicting conclusions.

GenAI feedback

GenAI feedback may significantly improve educational outcomes compared to traditional methods (Dann et al., 2024). Featuring comprehension, timeliness, and objectivity, GenAI may provide accurate feedback on students’ learning progress based on big data analyses and algorithms (Nash, 2024). GenAI may provide detailed feedback for students to find out their problems and improve their learning outcomes. Human teachers may fail to provide timely and comprehensive feedback due to the limitations of individual time, energy, and cognition. For instance, teachers may not provide detailed and in-depth feedback on each question (Lin and Crosthwaite, 2024), leading to delayed guidance. Therefore, GenAI feedback outperforms human feedback in the improvement of educational outcomes.

Nevertheless, conflicting results were also reported. For instance, no significant differences were reported in the effect of GenAI feedback and non-GenAI feedback in terms of interactive and negative emotions. No significant differences were also revealed in cognitive loads and visual information processing (Cosentino et al., 2025). Meanwhile, Henderson et al. (2025) found that students’ confidence in GenAI feedback was lower than teachers’. Cosentino et al. (2025) argued that the cognitive load was caused due to GenAI’s lack of emotional response, which was not addressed by the meta-cognition theory. GenAI’s inability to recognize emotions caused students’ lower trust in GenAI feedback (Henderson et al., 2025), which was inconsistent with sociocultural learning theory (Vygotsky, 1978), focusing on context-rich interaction. Until present, it is still hard for GenAI to completely understand human emotions and expressions, leading to the fact that GenAI may not outperform human teachers in academic achievements related to emotions and expressions.

Game-assisted GenAI

Although integrating games or gamification into GenAI may enhance students’ learning interest, interactions, activities, and engagement, game-assisted GenAI may not greatly improve educational outcomes compared with traditional methods. Games may have distracted students (Zehner and Hahnel, 2023) and urged them to concentrate on games rather than learning activities. Non-game-assisted GenAI focuses more on knowledge delivery and skill training, which may provide direct and systematic learning support. Educators and policymakers should pay attention to the balance between games and learning (Chien et al., 2024). This is in line with the self-determination theory (Chang, 2026), which could satisfy learners’ needs for competence improvement and autonomy enhancement. This contrast indicates that game design, rather than gamification itself, could greatly influence the e-learning effectiveness in game-assisted GenAI. Without the disturbance of games, students may focus on learning activities. Therefore, it is reasonable to conclude that game-assisted GenAI may not cause significantly higher educational outcomes than non-game-assisted GenAI.

However, researchers also revealed inconsistent findings. Chien et al. (2024) found that game-assisted design could increase learning engagement by providing rewarding systems and introducing competition mechanisms, which could arouse students’ learning interest and stimulate their enthusiasm, leading to improved learning effectiveness. Besides, if the game is properly integrated into GenAI technologies, catering to educational goals, the game-assisted GenAI instruction will yield positive educational outcomes. In addition, age and learning features could also influence the educational outcomes. Compared with older learners, younger learners may be more vulnerable to gamification addiction and digital distraction due to their weaker self-regulation, leading to a negative impact on educational outcomes. Therefore, for younger learners, non-gamed design may be a better choice.

China and Pakistan

In both China and Pakistan, GenAI may greatly improve educational outcomes compared with traditional methods, which may be enhanced by economic cooperation (Naz et al., 2019). This may be due to the similar educational environments and needs in both countries. Both countries are developing educational information and accepting innovative educational technologies. Educational resources in both countries are not evenly distributed, which may be complemented by the application of GenAI technologies. GenAI technologies may enhance their learning experiences by providing rich academic resources for students (Almuhanna, 2024). Cultures and educational backgrounds may also exert a great influence on GenAI-assisted educational outcomes. GenAI technologies may design customized learning and teaching methods according to local cultures and educational traditions, meeting the diverse needs of different students. For instance, educators may integrate local cultural factors into course designs, enabling intimate relationships and learning friendliness.

The introduction of GenAI technologies aligns with the educational development of China and Pakistan (Ahmad et al., 2023). There are issues of unequal distribution of educational resources in both countries. Students have no convenient access to high-quality educational resources and timely teacher guidance in some remote areas in both countries. With powerful data processing and personalized learning promotion, GenAI technologies may provide rich learning resources and personalized guidance for those in remote places, bridging the gap in educational resources.

Under cultural and educational backgrounds, both countries underscore the importance of knowledge accumulation and skill training. GenAI technologies enable stakeholders to design instructional content and pedagogical approaches based on cultural features and educational traditions. For instance, GenAI may be integrated into local cultural contexts and experiential learning (Casillo et al., 2022). Students and teachers are strongly encouraged to exchange using GenAI technologies, further improving the application of GenAI in education.

Korea and Turkey

In Korea and Turkey, GenAI may not greatly improve educational outcomes compared with traditional methods. The educational systems and cultural backgrounds may make it hard to employ GenAI technologies in education. Their education may value traditional pedagogical approaches and teacher-student interactions. They may spend less time and effort accepting and employing new technologies. Furthermore, their educational systems may not be adapted to the use of GenAI technologies. For instance, course design and teaching evaluation may not be compatible with GenAI technologies (Hong et al., 2025). Meanwhile, local cultures and social environments may shape barriers to the development and application of GenAI technologies. Admittedly, the included studies are limited, which may have caused biased conclusions. There may be many more studies on the effect of GenAI technologies on educational outcomes since they may be written in Korean and Turkish, making them inaccessible to this study. Future research should include more studies to delve into the effects of GenAI in Korea and Turkey.

There are also limitations on the development, promotion, and application of GenAI technologies in both countries. For instance, the limitation on algorithms may make it hard to meet the various needs of students in both countries (Er-Rafyg et al., 2025). Different qualities of the training data may cause various qualities of learning resources and feedback. Besides, both countries vary in infrastructure, teacher training, and students’ preferences (Hwang et al., 2006), which makes it hard to provide unanimous technological support for all of the teachers and learners. Some students and teachers may be reluctant to adapt to the GenAI-assisted education, while others may be ready for the GenAI-assisted pedagogical approaches. To sum up, technological differences may reduce the effect of GenAI technologies on educational outcomes in both countries.

University and secondary educational levels

GenAI may greatly improve educational outcomes in both universities and secondary schools, but the strategies need to be tailored.

At the university level, GenAI technologies may provide rich academic resources and research tools (Y. H. Chen and Zhang, 2025). Lan et al. (2025) reported that university students who used GenAI technologies could complete the assignment faster in line with the “scholarly skill development” framework (Boyer, 1990). GenAI technologies may improve the academic achievement and motivation of middle school students (Blanca Ibanez et al., 2020).

Furthermore, university students tended to possess significantly higher AI literacy than secondary school students (Kong and Yang, 2025). Secondary schools shed more light on the application of GenAI to basic knowledge delivery and interest cultivation (Kumar and Sharma, 2025).

Conclusion

Major findings

Through the meta-analysis and systematic review, this study concludes that GenAI technologies may have the potential to improve academic achievements, higher-order thinking abilities, and writing skills compared with traditional technologies; GenAI feedback may cause significantly higher educational outcomes than non-GenAI feedback; Game-assisted GenAI may fail to cause significantly higher educational outcomes than non-game-assisted GenAI; GenAI technologies may cause significantly higher educational outcomes than the non-GenAI technologies in China and Pakistan, rather than in Korea and Turkey; GenAI technologies may cause significantly higher educational outcomes than non-GenAI technologies at both university and secondary levels.

Limitations

Although this study is rigorously designed following the PRISMA protocol, it is still limited in several aspects. Firstly, this study may not have included all the related studies due to the limitation of library resources, possibly leading to bias in the conclusion. Secondly, the heterogeneity in the included studies may have negatively influenced the meta-analytical results. Thirdly, GenAI technologies have been developing at a dramatic pace, which may have influenced educational outcomes continuously.

Future research directions

Future research should focus on the effect of GenAI technologies on educational outcomes in multifaceted manners. Future researchers should delve into the impact of GenAI technologies on students with disabilities (Adako et al., 2024) and design personalized teaching strategies meeting their needs. Meanwhile, future researchers should pay attention to the long-term influence of GenAI technologies on students’ self-directed learning motivation and higher-order thinking abilities (Borge et al., 2024). Finally, they should keep pace with the rapid development of new GenAI technologies and investigate the adaptive GenAI strategies in different countries with various cultures, realizing innovative educational development with the assistance of GenAI technologies and metaverse technologies (Chu et al., 2025).

Directions for future literature reviews

Future literature review studies should be further expanded in the following ways. Firstly, future research may extend the citations to five to eight high-quality peer-reviewed studies to investigate each theme. Secondly, the descriptive analysis may be replaced by analytical integration to examine how GenAI feedback reduces information processing loads (Gkintoni et al., 2025). Finally, researchers may establish an analysis framework ranging from technological integration to constructive learning activities (Cattaneo et al., 2024).