A comprehensive narrative review of artificial intelligence use in the diagnosis and management of metabolic dysfunction-associated steatotic liver disease
Introduction
Metabolic dysfunction-associated steatotic liver disease (MASLD) prevalence has risen by 50% over the last three decades, currently affecting nearly one-third of the global population (1). It exists as a spectrum of worsening severity, ranging from simple steatosis to metabolic dysfunction-associated steatohepatitis (MASH), fibrosis, cirrhosis, and hepatocellular carcinoma (HCC) (2,3) (Figure 1). Despite this escalating prevalence, the disease remains frequently underdiagnosed as the gold-standard for diagnosis—liver biopsy—is invasive, costly, and prone to sampling variability, making it unsuitable for mass screening (2). On the other hand, non-invasive tests (NITs), such as magnetic resonance elastography (MRE), have reasonable diagnostic accuracy, but the expense and limited availability prevent their widespread use (4). Standard diagnostic codes capture fewer than 1% of cases, according to a recent study (5). Consequently, the majority of the disease remains undetected. In contrast, opportunistic artificial intelligence (AI) screening of asymptomatic individuals identified steatosis in nearly half of the same population (6). Now, as we have moved towards the precise metabolic definition established by recent multi-society consensus (3), AI can help mine electronic health records (EHRs) (7), automatically detect the disease from routine imaging (6), and tailor personalized interventions (8), making it a valuable tool across the breadth of the MASLD management spectrum. This review aims to describe the current literature on the use and role of AI in the diagnosis and management of MASLD (Figure 1). Recent reviews have explored AI across various liver diseases. Our review focuses on the new MASLD nomenclature while including historical nonalcoholic fatty liver disease (NAFLD) data. Additionally, we explore advanced applications like multi-omics, digital twin technology, and clinical trial design. We present this article in accordance with the Narrative Review reporting checklist (available at https://tgh.amegroups.com/article/view/10.21037/tgh-2026-0002/rc).
Methods
For this review, we searched the PubMed/MEDLINE database for English-language articles published from January 2005 through December 2025. The search strategy was designed to capture studies utilizing various AI modalities across the clinical spectrum of MASLD from its risk prediction to histopathologist staging and personalized management. This included various AI modalities, including machine learning (ML), deep learning (DL), and natural language processing (NLP). We prioritized original research, randomized controlled trials, and meta-analyses. The detailed search strategy, including specific terms and inclusion/exclusion criteria, is summarized in Table 1.
Table 1
| Items | Specification |
|---|---|
| Date of search | December 15, 2025 |
| Databases searched | PubMed/MEDLINE |
| Search terms used | “Artificial Intelligence”, “Machine Learning”, “Deep Learning”, “MASLD”, “NAFLD”, “NASH”, “Fibrosis”, “Steatosis” |
| Timeframe | January 2005 to December 2025 |
| Inclusion and exclusion criteria | Inclusion: English-language original research, meta-analyses, review articles and clinical trials regarding AI in liver disease. Exclusion: abstracts without full text; case reports; non-English literature |
| Selection process | The literature search and initial article selection were conducted by the first author (A.N.). Final inclusion of references was determined based on the authors’ collective judgment regarding each article’s relevance, historical context, and significance to the clinical care continuum of MASLD |
AI, artificial intelligence; MASLD, metabolic dysfunction-associated steatotic liver disease; MEDLINE, Medical Literature Analysis and Retrieval System Online; NAFLD, nonalcoholic fatty liver disease; NASH, nonalcoholic steatohepatitis.
AI in epidemiology and risk prediction
As hospital systems increasingly adopt EHRs, AI algorithms can mine these large patient data repositories to identify patients with MASLD using tools like ML and NLP. Compared with traditional identification using diagnosis codes or standard imaging, AI methods demonstrate greater sensitivity (5,7,9). For instance, Rodriguez et al. revealed that standard International Classification of Diseases (ICD) diagnosis codes capture less than 1% of actual MASLD cases, thereby leaving the majority of cases undetected (5). In contrast, Stuart et al. demonstrated that the NLP-MASLD algorithm, which scans radiology text and laboratory values, has a positive predictive value (PPV) of >93% (10).
Beyond identification of MASLD, AI models have consistently outperformed traditional scoring systems such as the fibrosis-4 index (FIB-4) and the NAFLD fibrosis score (NFS) in predicting progression to advanced fibrosis (AF) (11-23) (Table 2). A study by Dabbah et al. trained an Extreme Gradient Boosting (XGBoost) model on routine EHR data from 618 MASLD patients, achieving an area under the curve (AUC) of 0.91 for detecting AF (F3–F4), surpassing the FIB-4 model’s AUC (0.78). In addition, this model achieved a negative predictive value (NPV) of 99%, making it a highly effective tool for ruling out the disease in primary care settings (11). Similarly, Xiong et al. developed an XGBoost model using four simple lab indicators—triglycerides (TG), albumin (ALB), international normalized ratio (INR), and high-density lipoprotein (HDL)—that predicted AF with an AUC of 0.92, surpassing FIB-4 assessment (AUC of 0.75) (23). Results were similar in low-prevalence populations; Blanes-Vidal’s study showed that the LiverAID S model achieved an AUC of 0.91 and NPV of ≥0.98 for ruling out significant liver stiffness (>8 kPa), outperforming standard blood-based indices (14).
Table 2
| Study (first author, year) | AI model tested | Input data source | Prediction target | AI model performance (AUC) | FIB-4 performance (AUC) | Key findings/notes |
|---|---|---|---|---|---|---|
| Dabbah, 2025 (11) | XGBoost | Clinical and lab features from EHRs | Advanced fibrosis (F3–F4) | 0.91 | 0.78 | Outperformed standard non-invasive tests in primary care |
| Alkhouri, 2026 (12) | ALADDIN-F2-VCTE | Routine lab parameters + VCTE | Significant fibrosis (≥F2) | 0.79 | 0.66 | Outperformed VCTE, FAST, and Agile-3 |
| Chang, 2023 (13) | RF | 17 demographic/clinical features | Advanced fibrosis (≥F3) | 0.89 | 0.82 | Higher accuracy for all fibrosis stages (≥F2, ≥F3, F4) compared to FibroScan and FIB-4 |
| Blanes-Vidal, 2022 (14) | LiverAID S | 8 clinical variables + 5 serum markers | Significant liver stiffness (>8 kPa) | 0.91 | 0.70 | High NPV (≥0.98) for ruling out fibrosis in primary care |
| Calès, 2025 (15) | FIB-12 | 11 blood markers + LSM | Advanced fibrosis (F≥3) | 0.91 | 0.76 | Outperformed all other non-invasive tests evaluated |
| Fan, 2024 (16) | LSM-plus model | LSM + 5 clinical indices (age, sex, PLT, ALB, TBil) | Advanced fibrosis (F≥3) | 0.88 | 0.80 | Higher accuracy than LSM alone and the Agile score |
| Feng, 2021 (17) | RF (MLA) | 5 variables (BMI, PC-III, IV-C, AST, A/G ratio) | Significant fibrosis (F≥2) | 0.89 | 0.58 | Superior to FIB-4, NFS, APRI |
| Kalka, 2025 (18) | Fibro-Predict (XGBoost) | Routine blood tests from EHRs | 5-year risk of cirrhosis & AF | 0.79 | 0.71 | Predicts future cirrhosis risk and detects current advanced fibrosis |
| Liu, 2025 (19) | XGBoost | Routine clinical and laboratory data | Significant fibrosis (F2–F4) | 0.71 | 0.54 | Outperformed seven existing non-invasive scores |
| Mamandipoor, 2023 (20) | XGBoost | Lab, clinical data, diet | Liver fibrosis (LSM >8 kPa) | 0.75 | 0.61 | Moderate accuracy; noted better performance in males |
| Njei, 2024 (21) | XGBoost | Demographic, clinical, lab data | High-risk MASH (FAST ≥0.35) | 0.95 | 0.50 | Outperformed FIB-4, APRI, BARD, and NFS |
| Sarvestany, 2022 (22) | Ensemble | Lab, clinical, demographic data | Advanced fibrosis (F3–F4) | 0.87 | 0.83 | Reliable performance across diverse validation cohorts |
| Xiong, 2025 (23) | XGBoost | 4 lab indicators (TG, ALB, INR, HDL) | Advanced fibrosis (F3/F4) | 0.92 | 0.75 | Higher AUC compared to other models in the study |
A/G, albumin-to-globulin ratio; AF, advanced fibrosis; AI, artificial intelligence; ALB, albumin; APRI, AST to platelet ratio index; AST, aspartate aminotransferase; AUC, area under the curve; BARD, body mass index, AST/ALT ratio, diabetes; BMI, body mass index; EHR, electronic health record; FAST, FibroScan-AST; FIB-4, fibrosis-4 index; HDL, high-density lipoprotein; INR, international normalized ratio; IV-C, type IV collagen; LSM, liver stiffness measurement; MASH, metabolic dysfunction-associated steatohepatitis; MASLD, metabolic dysfunction-associated steatotic liver disease; MLA, machine learning algorithm; NFS, NAFLD fibrosis score; NPV, negative predictive value; PC-III, procollagen type III; PLT, platelet; RF, random forest; TBil, total bilirubin; TG, triglycerides; VCTE, vibration-controlled transient elastography; XGBoost, Extreme Gradient Boosting.
Other approaches have focused on integrating specific biomarkers to enhance performance. The “FIB-12” model, which combines 11 blood markers with liver stiffness measurement (LSM), achieved an AUC of 0.91 for diagnosing AF, outperforming all other NITs evaluated (15). Furthermore, Chang et al. used a random forest model incorporating 17 demographic and clinical features, achieving an AUC of 0.89 for detecting AF, outperforming both FibroScan and FIB-4 (13). Recently, the ALADDIN study introduced a machine-learning-based web calculator that estimates the likelihood of significant fibrosis (≥F2). In external validation, this model achieved an AUC of 0.79, outperforming vibration-controlled transient elastography (VCTE) alone (0.75), thereby offering an alternative for risk stratification (12).
By integrating multi-omics data (genomics, proteomics, and metabolomics), AI enables a deeper, more personalized understanding of MASLD. The FibroGENE-DT model, which incorporates the IFNL4 genotype into a decision tree algorithm, resulted in an AUC of 0.80 for significant fibrosis prediction (24). Additionally, Stefanakis et al. developed an ML model using clinical and metabolomic variables to accurately detect MASH with fibrosis stage F2–F3 (the specific group eligible for resmetirom therapy in their study). This model achieved an AUC of 0.94, outperforming other NITs (AUC 0.59–0.76) and highlighting the potential of AI to guide therapeutics (25) (Table 3).
Table 3
| Study (first author, year) | AI model tested | Input data source | Prediction target | AI model performance (AUC) | FIB-4 performance (AUC) | Key findings/notes |
|---|---|---|---|---|---|---|
| Eslam, 2016 (24) | FibroGENE-DT | 1 SNP + clinical/lab data | Significant fibrosis (≥F2) | 0.80 | 0.75 | Integrating a single genetic marker improved accuracy |
| Stefanakis, 2025 (25) | ML model (CatBoost) | Clinical variables + 2 metabolites | MASH with fibrosis F2–F3 | 0.94 | 0.59 | Predicts the specific MASH F2–F3 range for resmetirom eligibility. Outperformed other NITs |
| Huang, 2025 (26) | Machine Learning Score (RF) | Serum metabolites + clinical | MASH | 0.87 | Not reported | Predicts MASH and mortality better than FIB-4 or NFS |
| Luo, 2021 (27) | Elastic-Net Algorithm | SOMAscan proteomics | Advanced fibrosis (F3–4) | 0.78 | 0.74 | Performed similarly to or better than FIB-4, APRI and NFS |
AI, artificial intelligence; APRI, AST to platelet ratio index; AUC, area under the curve; DT, decision tree; FIB-4, fibrosis-4 index; MASH, metabolic dysfunction-associated steatohepatitis; MASLD, metabolic dysfunction-associated steatotic liver disease; ML, machine learning; NFS, NAFLD fibrosis score; NIT, non-invasive test; RF, random forest; SNP, single nucleotide polymorphism.
AI-enhanced imaging modalities
AI augments the ability of imaging studies to detect liver disease with higher diagnostic performance and less operator variability (Table 4). It is important to note that fibrosis imaging features and model performance can be etiology-dependent; hence, models trained on mixed etiology or chronic viral hepatitis cohorts may exhibit different diagnostic thresholds when applied exclusively to MASLD populations. Ultrasound (US) is the most frequently used first-line imaging tool to detect liver fat in MASLD, but it is often limited by inter-observer variability. AI mitigates this by employing DL models to standardize interpretation. For example, Yang et al. trained a two-stream convolutional neural network (2S-NNet) on abdominal US images, achieving an AUC of 0.90 for hepatic steatosis detection, surpassing the diagnostic performance of five conventional fatty liver indices (28). Similarly, Liu et al. developed the “Deep learning-based Integration of Dual-mode ultrasound and Clinical data” (DIDL) model, which integrates US contrast-enhanced micro-flow cines with B-mode images and clinical parameters. This multimodal approach achieved an AUC of 0.901 for diagnosing significant fibrosis, outperforming radiologist visual assessment (AUC 0.800) and FIB-4 scoring (AUC 0.776) (29). Furthermore, Ruan et al. developed the MSTNet (Multi-Scale Texture Network) model, which detected liver fibrosis with an AUC of 0.92, higher than the aspartate aminotransferase-to-platelet ratio index (APRI) (AUC 0.67) and FIB-4 (AUC 0.71) (30).
Table 4
| Study (first author, year) | AI model tested | Input data source | Prediction target | AI model performance (AUC) | FIB-4 performance (AUC) | Key findings/notes | Cohort etiology |
|---|---|---|---|---|---|---|---|
| Yang, 2023 (28) | 2S-NNet (CNN) | Abdominal ultrasound images | Hepatic steatosis | 0.9 | Not evaluated | Predicts steatosis, not fibrosis. Outperformed five fatty liver indices | MASLD cohort |
| Liu, 2023 (29) | DIDL model | Ultrasound + clinical data | Significant liver fibrosis (≥F2) | 0.901 | 0.776 | Integrating multimodal ultrasound data provided better diagnostic performance than APRI and FIB-4 | HBV-only |
| Ruan, 2021 (30) | MSTNet | B-mode ultrasound images | Significant fibrosis (≥F2) | 0.92 | 0.71 | Outperformed APRI, FIB-4, and sonographers’ visual assessment | HBV-only |
| Lu, 2020 (31) | FibroBox (ML model) | Lab, ultrasound, and TE/FibroScan | Significant fibrosis (≥F2) | 0.87 | 0.67 | Significantly higher AUCs compared to TE, APRI, and FIB-4 | HBV-only |
Fibrosis imaging features and AI model performance can be etiology-dependent. Diagnostic thresholds for models trained on HBV-only or mixed etiology cohorts may differ when applied exclusively to MASLD populations. AI, artificial intelligence; APRI, AST to platelet ratio index; AUC, area under the curve; CNN, convolutional neural network; FIB-4, fibrosis-4 index; HBV, hepatitis B virus; MASLD, metabolic dysfunction-associated steatotic liver disease; ML, machine learning; TE, transient elastography.
Computed tomography (CT) scans are performed for various indications and can be leveraged for opportunistic screening for MASLD (Table 5). Visual analysis of fibrosis on CT typically has lower sensitivity than AI algorithms. Pickhardt et al. developed a fully automated tool for identifying liver steatosis on contrast-enhanced CT that demonstrated high concordance with unenhanced CT, enabling population-level screening without additional radiation exposure (6). For fibrosis staging, Choi et al. developed a DL model using contrast-enhanced CT images that achieved an AUC of 0.96 for diagnosing significant fibrosis (≥F2), surpassing radiologists’ visual interpretation (AUC 0.84) (32). This utility extends beyond contrast-enhanced models; Tang et al. used a clinical-radiomic model with non-contrast scans to predict liver steatosis, achieving an AUC of 0.88 (35).
Table 5
| Study (first author, year) | AI model tested | Input data source | Prediction target | AI model performance (AUC) | FIB-4 performance (AUC) | Key findings/notes | Cohort etiology |
|---|---|---|---|---|---|---|---|
| Choi, 2018 (32) | Deep Learning System | Contrast-enhanced CT images | Significant fibrosis (≥F2) | 0.96 | 0.84 | Outperformed two experienced radiologists and showed higher accuracy for staging liver fibrosis across all stages | Mixed etiology |
| Cui, 2021 (33) | CatBoost | Multiphase CT radiomics | Staging liver fibrosis | 0.65–0.80 | 0.56–0.59 | Outperformed radiologists’ interpretation for significant and advanced LF | Mixed etiology |
| Wang, 2022 (34) | R-fibrosis | Contrast-enhanced CT + GGT, PLT, ALB | Advanced fibrosis (F3-F4) | 0.883 | 0.714 | Combined radiomics and clinical marker model showed good diagnostic performance for advanced fibrosis and cirrhosis | Mixed etiology |
Fibrosis imaging features and AI model performance can be etiology-dependent. Diagnostic thresholds for models trained on HBV-only or mixed etiology cohorts may differ when applied exclusively to MASLD populations. AI, artificial intelligence; ALB, albumin; AUC, area under the curve; CT, computed tomography; FIB-4, fibrosis-4 index; GGT, gamma-glutamyl transferase; HBV, hepatitis B virus; LF, liver fibrosis; MASLD, metabolic dysfunction-associated steatotic liver disease; PLT, platelet.
Magnetic resonance imaging (MRI) is the gold standard modality for non-invasive assessment, yet AI can further augment its diagnostic ability (Table 6). While MRE can accurately stage MASLD, it is resource-intensive and not widely available. To address this, the 3D fibrosis-aware deep learning (3D FADL) model combines standard MRI sequences with non-imaging data to achieve an AUC of 0.90 for detecting significant fibrosis, potentially reducing the need for specialized elastography hardware (36). Similarly, Zheng et al. demonstrated that a combined convolutional neural network (CNN) model—integrating standard MRI DL features and routine biomarkers—achieved an AUC of 0.86, superior to that of FIB-4 (AUC 0.69) (37).
Table 6
| Study (first author, year) | AI model tested | Input data source | Prediction target | AI model performance (AUC) | FIB-4 performance (AUC) | Key findings/notes | Cohort etiology |
|---|---|---|---|---|---|---|---|
| Li, 2025 (36) | Full model (3D FADL) | MRI sequences + non-image data | Significant liver fibrosis (≥F2) | 0.9 | Not reported | Integrating MRI and non-image data significantly improves performance | Mixed etiology |
| Zheng, 2024 (37) | Combined Model (CNN + biomarkers) | Liver MRI + serum biomarkers | Cirrhosis (F4) | 0.86 | 0.69 | Superior to clinical models including FIB-4 and APRI | Mixed etiology |
| Zha, 2024 (38) | CoRC model (DL + LR) | MRI + GGT, PLT, ALB | Significant liver fibrosis (≥F2) | 0.81 | 0.67 | Fully automated model accurately triaged patients. Outperformed APRI and FIB-4 | Mixed etiology |
Fibrosis imaging features and AI model performance can be etiology-dependent. Diagnostic thresholds for models trained on HBV-only or mixed etiology cohorts may differ when applied exclusively to MASLD populations. AI, artificial intelligence; ALB, albumin; APRI, AST to platelet ratio index; AUC, area under the curve; CNN, convolutional neural network; DL, deep learning; FADL, fibrosis-aware deep learning; FIB-4, fibrosis-4 index; GGT, gamma-glutamyl transferase; HBV, hepatitis B virus; LR, logistic regression; MASLD, metabolic dysfunction-associated steatotic liver disease; MRI, magnetic resonance imaging; PLT, platelet.
AI in histopathology
The gold standard for the diagnosis of MASLD remains histology, yet it is limited by significant inter-pathologist variability and the categorical nature of traditional scoring systems (Table 7). AI has the potential to standardize interpretation and identify subtle histologic changes that might be missed by visual assessment alone. Abdurrachim et al. demonstrated that an AI-based digital pathology platform significantly improved inter-pathologist agreement for fibrosis staging, raising the Kappa score from 0.63 (moderate agreement) to 0.73 (substantial agreement). Furthermore, AI assistance improved the agreement for identifying patients with significant fibrosis (F2/F3) from 45% to 71% (39).
Table 7
| Study (first author, year) | AI model tested | Input data source | Prediction target | AI model performance | FIB-4/pathologist performance | Key findings/notes |
|---|---|---|---|---|---|---|
| Abdurrachim, 2025 (39) | AI Digital Pathology (HistoIndex/Genesis 200) | Unstained SHG/TPEF images + H&E/MT | Inter-pathologist agreement for fibrosis staging | Kappa: 0.73 (AI-assisted) | Kappa: 0.63 (manual read) | AI assistance increased pathologist concordance for clinical trial inclusion (F2/F3) from 45% to 71% and reduced the need for adjudication by ~25% |
| Naoumov, 2024 (40) | qFibrosis & qSepta (SHG/TPEF) | Unstained liver biopsies (F3 stage) | Fibrosis dynamics (progression vs. regression) | Detection rate: 82% (detected change in 14/17 placebo patients) | Detection rate: 36% (conventional staging found change in only 6/17 placebo patients; 64% were “no change”) | AI “Radar Maps” revealed zonal fibrosis changes (e.g., perisinusoidal) that conventional staging missed. Successfully differentiated progressive vs. regressive septa (P<0.001) |
| Ratziu, 2024 (41) | PathAI ML Models | Digitized H&E and Masson’s trichrome slides | Treatment response (fibrosis reduction) | P value: 0.0099 (ML continuous score detected reduction with semaglutide 0.4 mg) | P value: not significant (pathologist categorical score did not reach significance for fibrosis improvement alone) | ML continuous scores detected a statistically significant antifibrotic effect of semaglutide 0.4 mg that was missed by conventional pathologist staging |
| Sanyal, 2022 (42) | PathAI ML Fibrosis Model | Digitized Masson’s trichrome-stained slides | Liver-related clinical events (decompensation/death) | HR: 1.18 (95% CI: 1.08–1.28) per 0.1-unit increase | HR: 1.34 (95% CI: 1.26–1.44) per 1 unit increase (FIB-4) | Higher baseline ML fibrosis scores were significantly associated with an increased risk of clinical events (HR 1.18). Change in ML score also predicted disease progression (HR 1.12) |
AI, artificial intelligence; CI, confidence interval; FIB-4, fibrosis-4 index; H&E, hematoxylin and eosin; HR, hazard ratio; Kappa, Cohen’s kappa coefficient; ML, machine learning; MT, Masson’s trichrome; qFibrosis, quantitative fibrosis; SHG, second harmonic generation; TPEF, two-photon excitation fluorescence.
Beyond standardization, AI offers superior sensitivity for detecting disease progression or regression. Naoumov et al. utilized qFibrosis to assess collagen features and found that AI detected subtle fibrosis changes in 82% of placebo patients, whereas conventional staging identified changes in only 36% (40). Furthermore, AI can detect therapeutic efficacy that manual staging fails to capture. Ratziu et al. reported that, while conventional pathologist scoring did not show a significant reduction in fibrosis with semaglutide 0.4 mg, AI-based continuous scoring detected a statistically significant improvement (P=0.0099) (41). Additionally, Sanyal et al. showed that higher baseline ML scores are associated with an increased risk of liver-related clinical events [hazard ratio (HR) 1.18 per 0.1-unit increase], validating their prognostic utility (42).
AI in drug development and clinical trials
In addition to prediction, detection, and prognostication, AI can also play an important role in drug development (40,41,43-45) (Table 8). It can identify optimal candidates for a specific therapy (enrichment) and detect treatment responses that conventional methods miss. For enrichment, Feng et al. used ML to identify novel biomarkers (such as CNPY4 and ENTPD6) associated with MASLD and HCC, achieving superior diagnostic capability (AUC 0.941) (46).
Table 8
| Trial name (NCT ID) | Drug/mechanism | Phase | AI tool used | Key AI-derived outcome/insight |
|---|---|---|---|---|
| FLIGHT-FXR (NCT02855164) (40) | Tropifexor (FXR agonist) | Phase 2 | qFibrosis (HistoIndex) | The AI model was able to detect regression vs. progression in liver fibrosis |
| Semaglutide Phase 2 (NCT02970942) (41) | Semaglutide (GLP-1 receptor agonist) | Phase 2 | PathAI ML Models | AI continuous scoring detected a statistically significant fibrosis reduction (P=0.0099) that could not be detected using conventional categorical staging |
| IMPACT (NCT05989711) (43) | Pemvidutide (GLP-1/Glucagon agonist) | Phase 2b | Liver Explore (PathAI) | AI detected “substantial decreases in the proportionate area of pathological fibrosis” (>50% reduction in 35% of high-dose patients), providing granular efficacy data beyond standard staging |
| NN9931-4492 (NCT03987451) (44) | Semaglutide (GLP-1 receptor agonist) | Phase 2 | PathAI ML Software | AI agreed with pathologists that there was no significant fibrosis improvement vs. placebo |
| ATLAS (NCT03449446) (45) | Cilofexor/Firsocostat | Phase 2b | Machine Learning (PathAI) | ML detected a significant shift in fibrosis patterns from F3–F4 to ≤F2 with combination therapy (P=0.040), that was missed with manual staging, suggesting potential therapeutic role |
AI, artificial intelligence; FXR, farnesoid X receptor; GLP-1, glucagon-like peptide-1; MASH, metabolic dysfunction-associated steatohepatitis; ML, machine learning; NCT, national clinical trial; qFibrosis, quantitative fibrosis.
Another important application of AI lies in optimizing clinical trial endpoints. Taylor-Weiner et al. developed the Deep Learning Treatment Assessment (DELTA) liver fibrosis score, which quantifies fibrosis on a continuous scale rather than in categorical stages. This AI-based score detected antifibrotic treatment effects that were missed by standard Nonalcoholic Steatohepatitis Clinical Research Network (NASH CRN) staging, effectively capturing therapeutic signals that manual pathology lacked the sensitivity to detect (47). Similarly, in the FLIGHT-FXR trial, Naoumov et al. employed AI-based second harmonic generation (SHG) microscopy to re-evaluate patients labeled as stable (no change) by pathologists. The AI analysis revealed that 82% of these patients had quantifiable fibrosis progression or regression, demonstrating that AI can detect granular changes in liver architecture that remain invisible to the human eye (40).
AI in clinical practice
The next critical step is the integration of AI into day-to-day clinical workflows and EHRs (Table 9). EHRs contain vast amounts of unstructured data that AI can mine to uncover undiagnosed MASLD. For instance, Stuart et al. developed an NLP algorithm that identified patients with hepatic steatosis, with a PPV of >93%, revealing a significant care gap: 85.4% of these patients lacked a formal clinical diagnosis (10). Additionally, AI can automate risk stratification; a systematic review by Njei et al. found that AI models could detect significant fibrosis (≥F2) with AUCs up to 0.95, potentially triggering automated alerts for hepatology referral (48).
Table 9
| Application type | AI tool/approach | Key study/evidence | Clinical function & outcome |
|---|---|---|---|
| EHR screening | NLP-MASLD Algorithm | Stuart et al., 2025 (10) | Automated case finding: scans radiology text and labs. Identified undiagnosed MASLD patients PPV >93% |
| Risk stratification | AI Risk Prediction Models | Njei et al., 2025 (48) | Automated triage: systematic review of AI models using routine clinical data achieved high accuracy (AUCs up to 0.94) for detecting significant fibrosis (≥F2) necessary to create automated risk stratification and referral systems |
| Personalized care | Digital Twin Technology | Joshi et al., 2023 (49) | Nutritional intervention: randomized trial of AI-guided personalized nutrition. Resulted in significant reduction in liver fat compared to standard care |
| Patient support | AI Chatbots (LLMs) | Pugliese et al., 2024 (50) | Remote counseling: validated LLMs (e.g., ChatGPT) provided accurate and safe answers to patient queries about MASLD management, scoring high on comprehensibility |
| Remote monitoring | AI-Integrated Wearables | Vemulapalli et al., 2025 (51) | Lifestyle tracking: AI analysis of wearable data enables continuous monitoring of physical activity and metabolic health, facilitating timely clinical interventions |
AI, artificial intelligence; AUC, area under the curve; EHR, electronic health record; LLM, large language model; MASLD, metabolic dysfunction-associated steatotic liver disease; NLP, natural language processing; PPV, positive predictive value.
Beyond screening, AI can transform individualized management through “Digital Twin” technology—a virtual replica of a patient’s physiology. In a randomized controlled trial, Joshi et al. demonstrated that a digital twin-guided nutrition program achieved a 72.7% remission rate from diabetes and a significant reduction in liver fat compared with standard care (49). AI’s utility also extends to digital therapeutics; Sato et al. reported that a smartphone-based “NASH App” intervention led to an average weight loss of 8.3%, with 68% of patients achieving histological improvement in disease activity (52). Furthermore, AI is being explored to guide personalized treatment decisions. By analyzing patient-specific multi-omic profiles and historical clinical trajectories, AI can help predict individual responses to emerging MASH therapies (such as resmetirom), thereby optimizing drug selection, minimizing adverse effects, and improving overall therapeutic efficacy (25). Finally, large language models (LLMs) can facilitate patient education. Pugliese et al. assessed ChatGPT as a counseling tool, concluding that it provided comprehensible answers and offered 24/7 support to complement physician guidance (50).
Challenges and limitations
While AI has the potential to revolutionize MASLD care, barriers must be addressed before it can be implemented in day-to-day clinical practice. A significant limitation is data heterogeneity and bias. Most (over 50 percent) of the AI studies in hepatology originate from just two countries (the United States and China), and algorithms are often developed using data without any patient demographic information, such as race or ethnicity (53,54). Additionally, most AI models are trained on retrospective data from large academic centers, which potentially limits generalizability and the prediction of future outcomes (20,55). Furthermore, the black-box nature of AI algorithms, in which a physician does not have a complete understanding of the complex internal workings of the AI and is unable to explain the way in which the diagnosis or risk prediction was obtained, creates a transparency gap that can lead to mistrust and potential liability issues should the AI miss a diagnosis (56,57). Finally, privacy laws restrict access to large datasets from multiple institutions for AI training, thereby continuing to limit the universal application of AI systems (20,55,57).
Finally, the deployment of AI in hepatology introduces substantial ethical challenges. Opportunistic screening complicates the informed consent process, as patients may incidentally be diagnosed with advanced liver disease. Furthermore, models trained on specific demographics can create algorithmic bias that risks worsening healthcare disparities if these tools yield suboptimal results for underrepresented minority populations.
Future directions
The future of MASLD care will be characterized by collaborative, secure, and personalized management. The optimization of AI algorithms requires adopting methods such as federated learning (FL) and digital twin technology, which enable models to be trained securely and globally across diverse healthcare institutions without sharing raw patient data, thereby mitigating data heterogeneity and bias (49,55,57,58) (Figure 2). Moreover, these AI advancements should also include multiomics and biomarker data to improve diagnostic and therapeutic capabilities (24-27,43,57,59-61).
Conclusions
The management of MASLD is positioned for a transformation with the adoption of AI. It will complement the role of clinicians by processing vast amounts of data to generate recommendations tailored specifically to individual patients. The technology has the potential to address the major challenges in MASLD management—underdiagnosis, staging variability, and ineffective lifestyle interventions. Moreover, the integration of AI into EHRs will make MASLD care more standardized and efficient. However, before this potential can be fully harnessed, concerted efforts must be made to address data heterogeneity, algorithmic bias, the black-box nature of AI algorithms, and privacy and security concerns. Addressing these limitations through secure, collaborative learning and explainable design will pave the way for AI to transition from a research novelty to an indispensable standard of care in MASLD management.
Acknowledgments
None.
Footnote
Reporting Checklist: The authors have completed the Narrative Review reporting checklist. Available at https://tgh.amegroups.com/article/view/10.21037/tgh-2026-0002/rc
Peer Review File: Available at https://tgh.amegroups.com/article/view/10.21037/tgh-2026-0002/prf
Funding: None.
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://tgh.amegroups.com/article/view/10.21037/tgh-2026-0002/coif). A.S. serves as an unpaid editorial board member of Translational Gastroenterology and Hepatology from August 2025 to July 2027. The other authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Younossi ZM, Golabi P, Paik JM, et al. The global epidemiology of nonalcoholic fatty liver disease (NAFLD) and nonalcoholic steatohepatitis (NASH): a systematic review. Hepatology 2023;77:1335-47. [Crossref] [PubMed]
- Castera L, Friedrich-Rust M, Loomba R. Noninvasive Assessment of Liver Disease in Patients With Nonalcoholic Fatty Liver Disease. Gastroenterology 2019;156:1264-1281.e4. [Crossref] [PubMed]
- Rinella ME, Lazarus JV, Ratziu V, et al. A multisociety Delphi consensus statement on new fatty liver disease nomenclature. Hepatology 2023;78:1966-86. [Crossref] [PubMed]
- Singh S, Venkatesh SK, Loomba R, et al. Magnetic resonance elastography for staging liver fibrosis in non-alcoholic fatty liver disease: a diagnostic accuracy systematic review and individual participant data pooled analysis. Eur Radiol 2016;26:1431-40. [Crossref] [PubMed]
- Rodriguez LA, Tucker LS, Saxena V, et al. Discrepancy in Metabolic Dysfunction-Associated Steatotic Liver Disease Prevalence in a Large Northern California Cohort. Gastro Hep Adv 2025;4:100630. [Crossref] [PubMed]
- Pickhardt PJ, Blake GM, Graffy PM, et al. Liver Steatosis Categorization on Contrast-Enhanced CT Using a Fully Automated Deep Learning Volumetric Segmentation Tool: Evaluation in 1204 Healthy Adults Using Unenhanced CT as a Reference Standard. AJR Am J Roentgenol 2021;217:359-67. [Crossref] [PubMed]
- Torgersen J, Skanderson M, Kidwai-Khan F, et al. Identification of hepatic steatosis among persons with and without HIV using natural language processing. Hepatol Commun 2024;8:e0468. [Crossref] [PubMed]
- Tincopa MA, Patel N, Shahab A, et al. Implementation of a randomized mobile-technology lifestyle program in individuals with nonalcoholic fatty liver disease. Sci Rep 2024;14:7452. [Crossref] [PubMed]
- Weng S, Hu D, Chen J, et al. Prediction of Fatty Liver Disease in a Chinese Population Using Machine-Learning Algorithms. Diagnostics (Basel) 2023;13:1168. [Crossref] [PubMed]
- Stuart A, Dotson A, Swarts K, et al. Development of an AI algorithm for early identification of MASLD in the electronic health record. Hepatol Commun 2025;9:e0834. [Crossref] [PubMed]
- Dabbah S, Mishani I, Davidov Y, et al. Implementation of Machine Learning Algorithms to Screen for Advanced Liver Fibrosis in Metabolic Dysfunction-Associated Steatotic Liver Disease: An In-Depth Explanatory Analysis. Digestion 2025;106:189-202. [Crossref] [PubMed]
- Alkhouri N, Cheuk-Fung Yip T, Castera L, et al. ALADDIN: A Machine Learning Approach to Enhance the Prediction of Significant Fibrosis or Higher in Metabolic Dysfunction-Associated Steatotic Liver Disease. Am J Gastroenterol 2026;121:362-74. [Crossref] [PubMed]
- Chang D, Truong E, Mena EA, et al. Machine learning models are superior to noninvasive tests in identifying clinically significant stages of NAFLD and NAFLD-related cirrhosis. Hepatology 2023;77:546-57. [Crossref] [PubMed]
- Blanes-Vidal V, Lindvig KP, Thiele M, et al. Artificial intelligence outperforms standard blood-based scores in identifying liver fibrosis patients in primary care. Sci Rep 2022;12:2914. [Crossref] [PubMed]
- Calès P, Canivet CM, Costentin C, et al. A new generation of non-invasive tests of liver fibrosis with improved accuracy in MASLD. J Hepatol 2025;82:794-804. [Crossref] [PubMed]
- Fan R, Yu N, Li G, et al. Machine-learning model comprising five clinical indices and liver stiffness measurement can accurately identify MASLD-related liver fibrosis. Liver Int 2024;44:749-59. [Crossref] [PubMed]
- Feng G, Zheng KI, Li YY, et al. Machine learning algorithm outperforms fibrosis markers in predicting significant fibrosis in biopsy-confirmed NAFLD. J Hepatobiliary Pancreat Sci 2021;28:593-603. [Crossref] [PubMed]
- Kalka IN, Hazzan R, Yacovzada NS, et al. Fibro predict a machine learning risk score for advanced liver fibrosis in the general population using Israeli electronic health records. Sci Rep 2025;15:32035. [Crossref] [PubMed]
- Liu M, Jiang L, Yang J, et al. Development and Validation of a Machine Learning-based Model for Prediction of Liver Fibrosis and MASH. J Clin Gastroenterol 2025; Epub ahead of print. [Crossref]
- Mamandipoor B, Wernly S, Semmler G, et al. Machine learning models predict liver steatosis but not liver fibrosis in a prospective cohort study. Clin Res Hepatol Gastroenterol 2023;47:102181. [Crossref] [PubMed]
- Njei B, Osta E, Njei N, et al. An explainable machine learning model for prediction of high-risk nonalcoholic steatohepatitis. Sci Rep 2024;14:8589. [Crossref] [PubMed]
- Sarvestany SS, Kwong JC, Azhie A, et al. Development and validation of an ensemble machine learning framework for detection of all-cause advanced hepatic fibrosis: a retrospective cohort study. Lancet Digit Health 2022;4:e188-99. [Crossref] [PubMed]
- Xiong FX, Sun L, Zhang XJ, et al. Machine learning-based models for advanced fibrosis in non-alcoholic steatohepatitis patients: A cohort study. World J Gastroenterol 2025;31:101383. [Crossref] [PubMed]
- Eslam M, Hashem AM, Romero-Gomez M, et al. FibroGENE: A gene-based model for staging liver fibrosis. J Hepatol 2016;64:390-8. [Crossref] [PubMed]
- Stefanakis K, Mingrone G, George J, et al. Accurate non-invasive detection of MASH with fibrosis F2-F3 using a lightweight machine learning model with minimal clinical and metabolomic variables. Metabolism 2025;163:156082. [Crossref] [PubMed]
- Huang Q, Qadri SF, Bian H, et al. A metabolome-derived score predicts metabolic dysfunction-associated steatohepatitis and mortality from liver disease. J Hepatol 2025;82:781-93. [Crossref] [PubMed]
- Luo Y, Wadhawan S, Greenfield A, et al. SOMAscan Proteomics Identifies Serum Biomarkers Associated With Liver Fibrosis in Patients With NASH. Hepatol Commun 2021;5:760-73. [Crossref] [PubMed]
- Yang Y, Liu J, Sun C, et al. Nonalcoholic fatty liver disease (NAFLD) detection and deep learning in a Chinese community-based population. Eur Radiol 2023;33:5894-906. [Crossref] [PubMed]
- Liu Z, Li W, Zhu Z, et al. A deep learning model with data integration of ultrasound contrast-enhanced micro-flow cines, B-mode images, and clinical parameters for diagnosing significant liver fibrosis in patients with chronic hepatitis B. Eur Radiol 2023;33:5871-81. [Crossref] [PubMed]
- Ruan D, Shi Y, Jin L, et al. An ultrasound image-based deep multi-scale texture network for liver fibrosis grading in patients with chronic HBV infection. Liver Int 2021;41:2440-54. [Crossref] [PubMed]
- Lu XJ, Yang XJ, Sun JY, et al. FibroBox: a novel noninvasive tool for predicting significant liver fibrosis and cirrhosis in HBV infected patients. Biomark Res 2020;8:48. [Crossref] [PubMed]
- Choi KJ, Jang JK, Lee SS, et al. Development and Validation of a Deep Learning System for Staging Liver Fibrosis by Using Contrast Agent-enhanced CT Images in the Liver. Radiology 2018;289:688-97. [Crossref] [PubMed]
- Cui E, Long W, Wu J, et al. Predicting the stages of liver fibrosis with multiphase CT radiomics based on volumetric features. Abdom Radiol (NY) 2021;46:3866-76. [Crossref] [PubMed]
- Wang J, Tang S, Mao Y, et al. Radiomics analysis of contrast-enhanced CT for staging liver fibrosis: an update for image biomarker. Hepatol Int 2022;16:627-39. [Crossref] [PubMed]
- Tang S, Wu J, Xu S, et al. Clinical-radiomic analysis for non-invasive prediction of liver steatosis on non-contrast CT: A pilot study. Front Genet 2023;14:1071085. [Crossref] [PubMed]
- Li W, Zhu Y, Zhao G, et al. Deep learning-based automated assessment of hepatic fibrosis via magnetic resonance images and nonimage data. Quant Imaging Med Surg 2025;15:8250-64. [Crossref] [PubMed]
- Zheng T, Zhu Y, Chen Y, et al. Fully automated MRI-based convolutional neural network for noninvasive diagnosis of cirrhosis. Insights Imaging 2024;15:298. [Crossref] [PubMed]
- Zha JH, Xia TY, Chen ZY, et al. Fully automated hybrid approach on conventional MRI for triaging clinically significant liver fibrosis: A multi-center cohort study. J Med Virol 2024;96:e29882. [Crossref] [PubMed]
- Abdurrachim D, Lek S, Ong CZL, et al. Utility of AI digital pathology as an aid for pathologists scoring fibrosis in MASH. J Hepatol 2025;82:898-908. [Crossref] [PubMed]
- Naoumov NV, Kleiner DE, Chng E, et al. Digital quantitation of bridging fibrosis and septa reveals changes in natural history and treatment not seen with conventional histology. Liver Int 2024;44:3214-28. [Crossref] [PubMed]
- Ratziu V, Francque S, Behling CA, et al. Artificial intelligence scoring of liver biopsies in a phase II trial of semaglutide in nonalcoholic steatohepatitis. Hepatology 2024;80:173-85. [Crossref] [PubMed]
- Sanyal AJ, Anstee QM, Trauner M, et al. Cirrhosis regression is associated with improved clinical outcomes in patients with nonalcoholic steatohepatitis. Hepatology 2022;75:1235-46. [Crossref] [PubMed]
- Noureddin M, Harrison SA, Loomba R, et al. Safety and efficacy of weekly pemvidutide versus placebo for metabolic dysfunction-associated steatohepatitis (IMPACT): 24-week results from a multicentre, randomised, double-blind, phase 2b study. Lancet 2025;406:2644-55. [Crossref] [PubMed]
- Loomba R, Abdelmalek MF, Armstrong MJ, et al. Semaglutide 2·4 mg once weekly in patients with non-alcoholic steatohepatitis-related cirrhosis: a randomised, placebo-controlled phase 2 trial. Lancet Gastroenterol Hepatol 2023;8:511-22. [Crossref] [PubMed]
- Loomba R, Noureddin M, Kowdley KV, et al. Combination Therapies Including Cilofexor and Firsocostat for Bridging Fibrosis and Cirrhosis Attributable to NASH. Hepatology 2021;73:625-43. [Crossref] [PubMed]
- Feng G, Targher G, Byrne CD, et al. Biomarker Discovery for Metabolic Dysfunction-associated Steatotic Liver Disease Utilizing Mendelian Randomization, Machine Learning, and External Validation. J Clin Transl Hepatol 2025;13:723-33. [Crossref] [PubMed]
- Taylor-Weiner A, Pokkalla H, Han L, et al. A Machine Learning Approach Enables Quantitative Measurement of Liver Histology and Disease Monitoring in NASH. Hepatology 2021;74:133-47. [Crossref] [PubMed]
- Njei B, Al-Ajlouni YA, Lemos SY, et al. AI-Based Models for Risk Prediction in MASLD: A Systematic Review. Dig Dis Sci 2026;71:1987-2014. [Crossref] [PubMed]
- Joshi S, Shamanna P, Dharmalingam M, et al. Digital Twin-Enabled Personalized Nutrition Improves Metabolic Dysfunction-Associated Fatty Liver Disease in Type 2 Diabetes: Results of a 1-Year Randomized Controlled Study. Endocr Pract 2023;29:960-70. [Crossref] [PubMed]
- Pugliese N, Polverini D, Lombardi R, et al. Evaluation of ChatGPT as a Counselling Tool for Italian-Speaking MASLD Patients: Assessment of Accuracy, Completeness and Comprehensibility. J Pers Med 2024;14:568. [Crossref] [PubMed]
- Vemulapalli B, Ghattu M, Atluri K, et al. Artificial Intelligence and Machine Learning Applications in Liver Disease. Clin Liver Dis 2025;29:755-70. [Crossref] [PubMed]
- Sato M, Akamatsu M, Shima T, et al. Impact of a Novel Digital Therapeutics System on Nonalcoholic Steatohepatitis: The NASH App Clinical Trial. Am J Gastroenterol 2023;118:1365-72. [Crossref] [PubMed]
- Celi LA, Cellini J, Charpignon ML, et al. Sources of bias in artificial intelligence that perpetuate healthcare disparities-A global review. PLOS Digit Health 2022;1:e0000022. [Crossref] [PubMed]
- Straw I, Wu H. Investigating for bias in healthcare algorithms: a sex-stratified analysis of supervised machine learning models in liver disease prediction. BMJ Health Care Inform 2022;29:e100457. [Crossref] [PubMed]
- Jagadeesh A, Aramrat C, Rai S, et al. Diagnostic accuracy of convolutional neural networks in classifying hepatic steatosis from B-mode ultrasound images: a systematic review with meta-analysis and novel validation in a community setting in Telangana, India. Lancet Reg Health Southeast Asia 2025;40:100644. [Crossref] [PubMed]
- Char DS, Shah NH, Magnus D. Implementing Machine Learning in Health Care - Addressing Ethical Challenges. N Engl J Med 2018;378:981-3. [Crossref] [PubMed]
- Balsano C, Burra P, Duvoux C, et al. Artificial Intelligence and liver: Opportunities and barriers. Dig Liver Dis 2023;55:1455-61. [Crossref] [PubMed]
- Rieke N, Hancox J, Li W, et al. The future of digital health with federated learning. NPJ Digit Med 2020;3:119. [Crossref] [PubMed]
- Ghosh S, Zhao X, Alim M, et al. Artificial intelligence applied to 'omics data in liver disease: towards a personalised approach for diagnosis, prognosis and treatment. Gut 2025;74:295-311. [Crossref] [PubMed]
- Haas ME, Pirruccello JP, Friedman SN, et al. Machine learning enables new insights into genetic contributions to liver fat accumulation. Cell Genom 2021;1:100066. [Crossref] [PubMed]
- Khusial RD, Cioffi CE, Caltharp SA, et al. Development of a Plasma Screening Panel for Pediatric Nonalcoholic Fatty Liver Disease Using Metabolomics. Hepatol Commun 2019;3:1311-21. [Crossref] [PubMed]
Cite this article as: Niazi A, Singh B, Singh C, Thakral N, Batta A, Sohal A. A comprehensive narrative review of artificial intelligence use in the diagnosis and management of metabolic dysfunction-associated steatotic liver disease. Transl Gastroenterol Hepatol 2026;11:68.

