1. Introduction
The contamination of water resources by emerging pollutants has become one of the most pressing environmental challenges of the 21st century. Unlike conventional contaminants such as heavy metals and organic dyes, emerging contaminants (ECs: pharmaceuticals, antibiotics, per- and polyfluoroalkyl substances (PFAS), endocrine-disrupting compounds (EDCs), microplastics, and pesticides) are continuously released into aquatic environments through wastewater discharge, agricultural runoff, and industrial effluents [1, 2]. These compounds are detected at trace concentrations (ng/L to μg/L), yet even at such levels they can disrupt endocrine function, promote antimicrobial resistance, and exert chronic toxicity on aquatic organisms [3, 4]. Conventional wastewater treatment plants are not designed to remove these micropollutants, and reported removal efficiencies vary widely depending on the compound, treatment configuration, and operating conditions [4].
Among the available remediation strategies, adsorption has emerged as one of the most versatile and cost-effective technologies for EC removal from aqueous systems [5, 6]. Its advantages include operational simplicity, high removal efficiency across a broad range of pollutant classes, and the availability of diverse adsorbent materials, from low-cost agricultural waste and biochar to engineered nanomaterials such as metal–organic frameworks (MOFs) and graphene-based composites [7]. However, adsorption performance depends on a complex interplay of factors, including adsorbent properties (surface area, pore structure, functional groups), solution chemistry (pH, temperature, ionic strength), and contaminant characteristics (molecular size, polarity, charge). The highly non-linear and multivariable nature of these interactions makes it difficult to predict adsorption behavior using conventional empirical models such as Langmuir, Freundlich, or pseudo-kinetic equations alone [6].
Machine learning (ML) offers a fundamentally different approach to
this modeling challenge. Rather than assuming a predefined functional
relationship, ML algorithms learn complex, non-linear mappings
directly from data, enabling them to capture interactions that elude
classical models [8]. In the broader field of
environmental engineering, ML has been successfully applied to water
quality prediction, process optimization, and pollutant fate modeling
[9]. For adsorption systems specifically, ML models have
been trained to predict adsorption capacity, removal efficiency, and
breakthrough curves from physicochemical descriptors, often achieving
coefficients of determination (
Despite this momentum, the literature on ML-assisted adsorption of ECs remains fragmented. Individual studies typically focus on a single adsorbent–contaminant system, employ different algorithms and validation strategies, and report inconsistent performance metrics, making it difficult to draw generalizable conclusions. The intersection of ML, adsorption and ECs has not gone unreviewed, and it is worth being exact about what the existing reviews do and do not provide, because a claim of novelty is only as good as the comparison behind it. Table 1 sets thirteen reviews published between 2022 and 2026 against the present work. Several are defined by adsorbent rather than contaminant, covering biochar [12], carbonaceous materials [13] or metal–organic frameworks [14]; several address conventional pollutants, principally heavy metals [15, 16, 17]; and others treat ML in adsorption generally without separating ECs as a class [18, 19, 20]. Three reviews published in 2026, after this review’s search cut-off of 24 February 2026 and therefore outside the studies it reviews, come closer. Fan et al. [21] address EC adsorption specifically, and Elngar et al. [22] report PRISMA adherence for adsorption-based treatment, while Yaghoobian et al. [23] cover PFAS across the whole management pipeline and compare interpretability strategies qualitatively. We therefore do not claim to be the first review of this subject.
What the comparison does show is a consistent difference in kind. Of the thirteen, eleven state neither the databases they searched nor the number of studies they included, and none reports pooling reported performance across studies quantitatively or synthesising feature-importance rankings across explainable-ML studies. Every DOI in Table 1 was checked against Crossref, and the verification records are provided in the Supplementary Information. The contribution of the present work is accordingly methodological rather than territorial: it is, to our knowledge, the first review of ML for EC adsorption to combine a stated and reproducible search and screening protocol with a quantitative meta-analysis of reported performance, a cross-study synthesis of SHAP-derived feature importance, and an assessment of reporting quality in the underlying literature. That this can be claimed at all is itself a comment on the field, since it is unusual for a body of reviews to leave the size and provenance of its evidence base unstated.
| Review | Pollutant scope | Databases searched | Studies included | Quant. meta-analysis | Cross-study XAI | Principal limitation relative to this work |
|---|---|---|---|---|---|---|
| Reynel-Avila et al. (2022) [18] | Organic and inorganic pollutants; ECs not a separate class | Web of Science |
|
No | No | ANN only; no protocol or eligibility criteria |
| Zhang et al. (2023) [19] | Water pollutants generally | Not stated | Not stated | No | No | No screening protocol; ECs not analysed separately |
| Zhang et al. (2023) [12] | Pollutants removed by biochar | Not stated | Not stated | No | No | Defined by adsorbent, not contaminant |
| Fiyadh et al. (2023) [15] | Heavy metals only | Not stated | Not stated | No | No | Single model family and single pollutant class |
| Yuan et al. (2023) [16] | Heavy metals, mainly on biochar | Not stated | Not stated | No | No | Conventional pollutants; no protocol |
| Wang et al. (2024) [13] | Organic pollutants on carbonaceous adsorbents | Not stated | 39 | No | No | Carbonaceous adsorbents only; cross-study statistics are algorithm tallies, not pooled effects |
| Sharmila et al. (2024) [14] | Organic pollutants on MOFs | Not stated | Not stated | No | No | Confined to MOFs; ML is one section of a materials review |
| El Mallahi et al. (2025) [24] | Not pollutant-specific (bibliometric) | Scopus | Not stated | No | No | Counts and clusters publications; extracts no performance data |
| Zhao et al. (2025) [17] | Heavy metals only | Not stated | Not stated | No | No | Conventional pollutants; feature importance not pooled |
| Yuan et al. (2025) [20] | Water-treatment adsorbents generally | Not stated | Not stated | No | No | Materials-design perspective; no defined set of studies |
| Fan et
al. (2026) |
Emerging contaminants | Not stated | Not stated | No | No | Same population as this work, but a decision-oriented framework rather than a synthesis over a defined set of studies; no protocol or study count |
| Yaghoobian et
al. (2026) |
PFAS only, whole management pipeline | Not stated | Not stated | No | Qualitative | Single EC class; interpretability compared narratively, not pooled |
| Elngar et
al. (2026) |
Dyes, heavy metals, organics | Not stated | Not stated | No | No | States PRISMA adherence but reports neither databases nor study count |
| This work | Emerging contaminants (PFAS, pharmaceuticals, EDCs, microplastics, ARGs, PCPs) | Scopus, PubMed (24 Feb 2026) | 202 | Yes | Yes (15 SHAP studies) | Two databases only; not prospectively registered |
To address this gap, we conducted a systematic review following the PRISMA 2020 guidelines, screening 2,164 records from Scopus and PubMed and analyzing a final set of 202 peer-reviewed studies that first appeared online between 2015 and 2025. Against the reviews in Table 1, which describe individual studies qualitatively, this work adds three analytical contributions: (i) a quantitative meta-analysis of algorithm performance using inverse-variance weighted pooling to generate cross-study comparisons; (ii) a cross-study synthesis of SHAP-based feature importance to identify universal and system-specific drivers of adsorption prediction; and (iii) a Minimum Reporting Checklist for Machine Learning in Adsorption Studies (MRCAS) and a technology readiness assessment for frontier ML architectures. We also report audits of our own screening and extraction procedures (Sections 2.6.1 and 2.6.2), on the view that a review which assesses the reporting quality of others should document its own. The specific objectives are: (1) to map the bibliometric landscape of ML-assisted adsorption research for ECs; (2) to identify which ML algorithms are most frequently applied and conduct a meta-analysis of their predictive performance; (3) to characterize the emerging contaminants and adsorbent materials studied and assess coverage gaps; (4) to evaluate methodological rigor, including dataset sizes, validation protocols, and reporting practices; (5) to synthesize feature importance findings across explainable ML studies; and (6) to provide a frontier technology roadmap and evidence-based recommendations for future work.
We restricted the scope to studies in which (1) ML is applied as a core modeling methodology, not merely cited; (2) adsorption is the primary removal mechanism; and (3) the target pollutant belongs to a recognized EC category. Studies addressing only conventional pollutants (heavy metals, dyes) without an EC component, or employing non-adsorption processes (photocatalysis, biodegradation, membrane filtration) as the primary mechanism, were excluded. This focused scope enables a rigorous and comparable assessment of the current state of ML in EC adsorption, while the lessons learned are broadly applicable to ML-assisted environmental modeling.
2. Methodology
This systematic review was conducted following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines. The completed checklist is provided in Table S1, and the subsections below are ordered to follow its Methods items so that each item maps to one subsection.
Parts of this review were automated. Screening was performed by rule-based scripts and one phase of data extraction by a language model. We report this at the level of detail required by PRISMA items 8 and 9, both of which ask for details of any automation tools used, and by the 2025 joint position statement on the use of artificial intelligence in evidence synthesis issued by Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence [25], which asks that such use be disclosed, that human oversight be described, and that the authors demonstrate rather than assert that methodological rigour was not compromised. The justification for automating is scale: 1,745 unique records and 383 full texts exceed what this author team could screen and extract by hand, and automation was adopted to make the review feasible rather than to make a feasible review faster. Because that choice substitutes reproducible rules for human judgement, we did not assume the substitution was safe. Both automated stages were audited against re-reading of the source records, and those audits, including the errors they found, are reported in Section 2.6.
2.1 Eligibility Criteria
Throughout this review, “machine learning” (ML) denotes the class of methods under study, and “artificial intelligence” (AI) is used only where it is part of a term taken from the source literature or the screening criteria, or where it refers to the language model used as an extraction tool in Section 2.5; the two are not used interchangeably. Studies were assessed against three core pillars for inclusion: (1) ML or AI methods must be applied as a core methodology, not merely cited; (2) adsorption must be the primary removal mechanism; and (3) the target pollutant must be an emerging contaminant. Table 2 summarizes the full inclusion and exclusion criteria.
| ID | Criterion |
|---|---|
| Inclusion criteria | |
| I1 | At least one ML algorithm applied to model or predict adsorption |
| I2 | Target contaminant belongs to an emerging category (PFAS, pharmaceuticals, EDCs, microplastics, ARGs, or PCPs) |
| I3 | Adsorption occurs in aqueous or liquid phase |
| I4 | Study reports quantitative performance
metrics (e.g., |
| I5 | Published between 2015 and 2025 |
| I6 | Written in English |
| I7 | Published in a peer-reviewed journal |
| Exclusion criteria | |
| E1 | Only conventional pollutants (heavy metals or dyes) without emerging contaminants |
| E2 | No ML/AI application (traditional statistics only) |
| E3 | Gas-phase adsorption only |
| E4 | Non-adsorption process as primary mechanism (photocatalysis, biodegradation, Fenton, etc.) |
| E5 | Non-peer-reviewed source (conference paper, preprint, book chapter) |
| E6 | ML mentioned only in future work, not applied |
| E7 | Duplicate publication across databases |
| E8 | Full text unavailable |
| E9 | Adsorption of pollutants onto microplastics (without microplastic removal) |
Review articles are excluded by criterion E5 as sources of primary data, but they were deliberately retained by the database filters described in Section 2.2, because the comparison of prior reviews in Table 1 and the backward citation checking both required them. Removal therefore happened downstream of retrieval rather than at the query, and Section 2.6.1 reports that this removal initially failed for two records.
2.2 Information Sources
Two electronic databases, Scopus and PubMed, were searched for articles published between January 2015 and December 2025. Results were restricted to peer-reviewed journal articles and reviews written in English. Both searches were executed on 24 February 2026, which is the cut-off date for this review: records indexed after that date were not retrieved, and no records were added subsequently. The publication window closes at 31 December 2025, so every article in the final set had been indexed for at least eight weeks before the search ran. A single dating convention is applied throughout: the year assigned to a study is the year it first appeared online, taken from the online publication date recorded in Crossref where one is available and from the date the DOI was registered where it is not. The issue label is not used. This matters for four records that a database’s indexed year would place outside the window. Three appeared online and were indexed during 2025 but carry a 2026 issue; one appeared online in August 2022 and was assigned to a 2024 issue almost two years later. Under the convention they are 2025 and 2022 literature respectively, and every study in the final set therefore falls within 2015–2025 with no exceptions. The four records and their Crossref evidence are listed in the dataset changelog accompanying the Supplementary Data. Two consequences follow and should be borne in mind when reading the temporal trends in Section 3.1. First, coverage of 2025 is likely to be marginally less complete than for earlier years, because indexing of late-2025 articles was still accruing when the search ran; the direction of the upward trend is robust to this, but the 2025 count should be read as a lower bound. Second, because the review is restricted to indexed, peer-reviewed literature, it inherits the publication bias of that literature, in which models that perform well are more likely to be reported than models that do not. The funnel plot in Section 3.5 addresses this directly.
This review was not registered in PROSPERO or another prospective register, and no protocol was published in advance; this is recorded in item 24 of the PRISMA checklist (Table S1) and is stated here because registration, while not mandatory, is the usual safeguard against post hoc changes to eligibility criteria. Two features of the workflow limit that risk. The eligibility criteria and the full Boolean query were fixed on 23 February 2026, one day before the searches were run, and neither was altered afterwards. And because every screening and extraction step was executed by scripts rather than by ad hoc judgement, the decision rules are recoverable in full from the code and the logged per-record decisions, which are available from the corresponding author on request.
2.3 Search Strategy
The search strategy employed a three-block Boolean query combining: (1) machine learning and artificial intelligence terms, (2) adsorption and sorption-related terms, and (3) emerging contaminant categories. The ML/AI block included terms such as “machine learning,” “artificial neural network,” “random forest,” “support vector machine,” “deep learning,” “gradient boosting,” and related algorithm names. The adsorption block covered “adsorption,” “biosorption,” “sorption capacity,” “isotherm,” and specific adsorbent materials. The emerging contaminant block encompassed six sub-categories: per- and polyfluoroalkyl substances (PFAS), pharmaceuticals and antibiotics, endocrine-disrupting compounds (EDCs), microplastics, antibiotic resistance genes (ARGs), and personal care products (PCPs), each with specific compound names and synonyms. The three blocks were combined using the AND operator. The full query strings as executed against each database are reproduced in SI Section S2.1.
The initial search yielded 1,337 records from Scopus and 827 from PubMed, totaling 2,164 records. After removing 419 duplicates (identified by DOI, PMID, and title matching), 1,745 unique records remained for screening.
2.4 Selection Process
The PRISMA flow diagram summarizing the study selection process is presented in Fig. 1.
2.4.1 Title and abstract screening
Title and abstract screening was performed using an automated dual-reviewer approach. It is worth being precise about what the two screeners were, because the term “AI-assisted” would overstate them: both are deterministic, rule-based scripts that apply curated regular-expression dictionaries with contextual weighting to the title, abstract and keyword fields. No language model was involved at any point in screening. Reviewer A used a regex-contextual scoring method and Reviewer B a complementary pattern-matching strategy built on independently compiled term lists; each returned INCLUDE, EXCLUDE or UNCERTAIN per record. Because the rules are fixed and the inputs are text fields, both screeners are deterministic: re-running them on the same records reproduces every decision exactly.
Decisions were combined by an explicit rule rather than by
discussion. Concordant decisions were carried through unchanged.
Where one screener returned UNCERTAIN, the confident screener
decided, in both directions (INCLUDE with UNCERTAIN yielded
inclusion; EXCLUDE with UNCERTAIN yielded exclusion). Direct
conflicts, in which one screener returned INCLUDE and the other
EXCLUDE, were not resolved by this rule but passed to a third
rule-based script (
A subsequent tightened screening phase applied stricter thresholds, requiring that ML/AI be demonstrably the core methodology, adsorption be the primary process, and the study be original research rather than a review. This phase excluded an additional 305 records (266 auto-removed, 39 borderline cases excluded after manual inspection), yielding 383 studies for full-text assessment.
2.4.2 Full-text screening
Full-text screening employed a section-aware scoring algorithm
that assigned different weights to matches found in different
parts of each paper. Matches in the Methods section received the
highest weight (
Of the 383 candidate papers, full texts were obtained for 382 (one paper could not be retrieved). Of these 382, 173 were directly included, 143 were excluded, and 66 were classified as borderline. The 66 borderline cases were manually reviewed by one author, resulting in 31 additional inclusions and 35 exclusions. The final set comprised 204 studies for data extraction and synthesis. Two of these were identified during revision as review articles rather than primary studies and were removed, giving the 202 studies analysed here (Section 2.6.1).
2.5 Data Collection Process and Data Items
A structured data extraction framework was developed to capture
over 30 fields from each included study, organized into five
categories: (1) bibliographic metadata (DOI, title, authors,
journal, year, country); (2) ML methodology (algorithms used,
best-performing model, validation method, dataset size);
(3) adsorption system (adsorbent material, target contaminant, EC
category); (4) model performance (
Three properties of the process should be stated before the phases are described, because each bounds what the extracted dataset can support. Data were collected by a single automated pass over each report, reconciled between two extraction methods but not duplicated by an independent second extractor. Study authors were not contacted to supply missing or ambiguous values. And no values were digitised from figures: extraction operated on text recovered from each PDF, so a quantity reported only inside a plot is systematically absent from the dataset. The verification in Section 2.6.2 measures the consequences of all three.
Full texts were extracted from the 383 candidate PDFs using the pypdf library, achieving a 92.1% good-quality text extraction rate. Data extraction was then performed in two complementary phases.
2.5.1 Phase 1: Regex-based extraction
An automated extraction pipeline applied 55 ML algorithm
patterns, 25+ adsorbent material categories, and 80+
contaminant-specific regular expressions to the extracted full
texts. This approach achieved high coverage for well-defined
fields such as ML algorithms (98%), EC category (99%), and country
(97%), but lower coverage for context-dependent fields requiring
interpretive reading, such as best-performing model (51%),
2.5.2 Phase 2: Language-model re-extraction
To address gaps in context-dependent fields, the same reports were re-extracted using a large language model. The model was Claude Opus 4.6 (Anthropic), accessed in February 2026 and operated through an interactive agent session rather than through a scripted API call. Three properties of that arrangement follow, and we state them rather than leave them to be inferred. The session used the provider’s default sampling settings, so the model is not deterministic: unlike the rule-based screeners of Section 2.4, re-running it is not guaranteed to reproduce every value. The operative instruction was not preserved verbatim, so the input specification and the output schema can be published (SI Section S2.4) but the prompt itself cannot. And no second, independent language-model pass was run. These three properties are the reason the accuracy of the output was measured directly rather than argued from the design (Section 2.6.2).
What the model was given, and what it was permitted to return,
are both recoverable and are reported in full. The 204 papers then
in the included set were prepared in 41 batches of 5 papers each.
For each paper, a structured text package was compiled from that
paper’s own PDF, containing the header (first 3,000 characters,
for affiliation information), abstract (up to 2,000 characters),
Methods section (up to 4,000 characters), Results and Discussion
section (up to 4,000 characters), and Conclusion (up to 2,000
characters). The model was given no access to any other source and
was not asked to recall the study from memory. Its output was
confined to a fixed schema of 11 context-dependent fields:
country, ML algorithms, best model, best
Results from both extraction phases were merged using a
systematic reconciliation protocol: when both sources provided
values, agreement was tracked; when only one source had a value,
it was retained; disagreements (
2.5.3 Standardization
A final standardization step applied canonical mappings to
seven key fields: country names, ML algorithm names, best model
identifiers, target contaminants, and adsorbent materials.
Algorithm abbreviations were unified (e.g., “LR,” “MLR”
2.6 Evaluation of Automation Tool Performance
Screening and extraction were both automated, so the reliability of this review rests on how well those two tools performed on this specific body of literature. Neither the agreement statistics of Section 2.4 nor the coverage gains of Section 2.5 answer that question: agreement measures whether two rule sets concur, and coverage measures how many cells were filled, while what matters is whether the decisions were right and the values correct. We therefore evaluated both tools against re-reading of the source records. Both evaluations were carried out by one of the authors, without an independent second assessor, which is a limitation we state plainly here and again in Section 3.10.
2.6.1 Audit of screening decisions
Because the merged rule can exclude a record on a single confident vote, the reliability of the included set depends on whether that rule discarded studies it should have kept. Agreement statistics cannot answer this; only re-reading the discarded records can. We therefore audited them directly. The 993 records excluded at the title/abstract stage were stratified by how the exclusion was reached (503 where one screener returned EXCLUDE and the other UNCERTAIN, 405 where both returned EXCLUDE, and 85 removed by the prescreening filter), and a proportionally allocated random sample of 101 records (10.2%, random seed 20260912) was drawn. Each sampled record was then re-assessed by an author against the same inclusion and exclusion criteria, from title and abstract, blind to the screeners’ stated reasons, and classified as one that should have proceeded to full text, one correctly excluded, or one the abstract cannot settle.
No record in the sample was judged a clear false exclusion (0 of 101; 95% Wilson CI 0–3.7%). Four records (4.0%) could not be settled from the abstract and would, on the audit rule, have proceeded to full text; taking all four as potential losses gives a conservative upper bound of 9.7% on the false-exclusion rate, or roughly 97 of the 993 exclusions. The four are informative individually. Two are review articles, which the tightening rule excludes on separate grounds; one models removal across whole treatment trains without naming an adsorption unit; and one models powdered-activated-carbon dosing for taste-and-odour compounds, which are not among the emerging-contaminant categories this review covers. The stratification bore out its premise only weakly: the one-sided stratum contributed three of the four unsettled records and the concordant stratum one, a difference well inside sampling error. The audit is reported by exclusion route in Table S12, with per-record judgements in SI Section S11.
The audit of exclusions prompted a complementary check of the
inclusions, and that check found a failure in the opposite
direction. Screening the titles of the included set against the
same review-article rule identified two records that are review
articles: one surveying machine learning for organic pollutant
adsorption on carbonaceous materials, which also appears in
Table 1 as prior
literature, and one reviewing predictive modelling of PFAS
behaviour. Their routes differ and both are informative. The first
was passed by both screeners as meeting the inclusion criteria.
The second was auto-included at the prescreening stage on keyword
strength alone, a route that bypasses the rule-based screeners and
therefore never met the review-article test. Both were removed,
which is why 204 studies leave screening and 202 enter the
synthesis. The effect on the reported results is confined to
descriptive counts: neither record carries an
Two further limitations of this audit should be stated. It samples the title/abstract stage only, so exclusions made at the tightening and full-text stages are not covered by it. And the assessor applied the same written criteria as the screeners, so it detects records wrongly removed under those criteria, not any narrowness in the criteria themselves.
2.6.2 Verification of extracted data
Coverage is not accuracy, and the gains reported in Section 2.5 say nothing about whether the values recovered are correct. Three properties of the pipeline constrain how far a language model could invent content: it was shown only text taken from the paper’s own PDF rather than being asked to recall the study, its output was confined to a fixed field schema, and every field was reconciled against the independent regex pass, with the 582 disagreements logged. None of these is a guarantee, so we measured the result rather than arguing from the design.
Thirty studies were drawn at random from the 198 whose full
text was available (random seed 20260912), and six fields that
carry the quantitative synthesis (best model, best-model
Where the pipeline recorded a value, it was right in 112 of 120
cells (93.3%, 95% CI 87.4–96.6%); active errors were 8 of all 180
cells (4.4%). Accuracy divides sharply by field type. The three
descriptive fields (best model, adsorbent material and target
contaminant) were correct in all 87 populated cells and were never
wrongly left empty. Every error and every omission falls among the
numeric and methodological fields, where accuracy when populated
was 75.8% and where 35 of 90 cells were wrongly empty. Validation
method is the weakest field by a wide margin: only half its
populated cells were correct, and 15 of 18 empty cells should have
carried a value. The errors are not random. Four of the eight are
the same failure, in which a three-way 60/20/20 partition was
compressed to “60:40 split”; two more attach a ratio from
elsewhere in the paper to the data split, once from an HPLC mobile
phase and once from an adsorbent activation recipe; and one
imports the
Two consequences follow, and we have applied both. First, the descriptive syntheses of algorithms, adsorbents and contaminants in Sections 3.2–3.4 rest on the fields the audit found reliable. Second, no claim in this review is now made from the absence of a value in a sparsely populated field, because the audit shows such absences are substantially extraction failures; where the argument requires knowing what studies actually reported, as for validation practice in Section 3.5, we rely on the 30 papers read in full rather than on field-occupancy counts.
2.7 Study Quality and Risk of Bias
No validated instrument exists for appraising risk of bias in machine-learning adsorption studies. Instruments developed for clinical prediction models, such as PROBAST, presuppose a diagnostic or prognostic outcome and a defined patient population, and they do not transfer to a regression of adsorption capacity on physicochemical descriptors. We therefore did not attempt a formal risk-of-bias appraisal of the included studies, and none should be read into Section 3.7: the eight-indicator score reported there measures what could be recovered from each study by automated extraction, which Section 2.6.2 shows is a lower bound on what the studies actually report, and it is a measure of machine-readability rather than of quality. The absence of a study-level appraisal is a limitation of this review and is recorded as one in Section 3.10.
Two related assessments were made and should not be confused with
the one that was not. Risk of bias arising from missing results
across the synthesis, that is publication bias, was assessed by
funnel plot and Egger’s regression and is reported in
Section 3.5. And
the methodological property that most threatens the pooled
estimates, namely that reported
Certainty in the body of evidence was not graded with GRADE or an equivalent framework. Those frameworks are built around an intervention-and-outcome question with a directional effect estimate, which this review does not pose: the pooled quantity here is a reported goodness-of-fit statistic, not a treatment effect, and its heterogeneity is reported directly instead (Section 3.5).
2.8 Effect Measures and Synthesis Methods
The effect measure carried into quantitative synthesis is the
Algorithm names were normalized to the canonical forms of Section 2.5 before pooling, so that a study reporting “GBR” and one reporting “GBDT” contribute to the same stratum. Pooled estimates are reported for the six most frequently reported best-performing algorithms.
Comparisons of
3. Results and Discussion
3.1 Bibliometric Overview
The 202 included studies, listed individually in Table S9, were
published between 2015 and 2025, with a clear upward trajectory
(Fig. 2a). Only
eight appeared before 2019, the earliest of them in 2015
[26], whereas the period from 2021 onward
accounted for 179 of the 202 studies (88.6%). The most productive
year was 2025 (
Geographically, research output was dominated by Asian
institutions
(Fig. 2b; country
assigned from the Scopus affiliation field of the first author).
China contributed the largest share of studies
(
The 202 studies were distributed across more than 90
peer-reviewed journals, indicating a broad interdisciplinary reach
(Fig. 2c). The most
frequent outlets were the Journal of Environmental Chemical
Engineering (
To characterize the broader research landscape beyond the 202 included studies, a bibliometric network analysis was performed on the full set of 1,337 Scopus-indexed records (Fig. 3). Country nodes were resolved from the last comma-separated element of each Scopus affiliation string. The keyword co-occurrence network (Fig. 3a) reveals distinct thematic clusters: a central core around “machine learning” and “adsorption” is tightly linked to “artificial neural network,” “random forest,” and “optimization,” while satellite clusters emerge around specific contaminant classes (“tetracycline,” “antibiotics,” “pharmaceuticals,” “microplastics,” “PFAS”) and adsorbent materials (“biochar,” “activated carbon”). The temporal overlay indicates that keywords such as “XGBoost,” “PFAS,” “microplastics,” “deep learning,” and “SHAP” (shown in lighter tones) represent the most recent research frontiers, consistent with the trends observed in the included studies. In contrast, earlier keywords such as “response surface methodology” and “genetic algorithm” appear in darker tones, reflecting their established but plateauing role.
The co-authorship country network (Fig. 3b) reveals that China and the United States serve as the two dominant hubs, with the strongest bilateral collaboration link between them. India, Iran, and South Korea form secondary hubs with extensive connections to both China and Western countries. European nations (United Kingdom, Germany, France, Spain) form a densely interconnected cluster, while countries in Southeast Asia, Africa, and Latin America appear as peripheral nodes with fewer collaboration links. This pattern underscores the geographic concentration noted earlier and highlights opportunities for expanding collaborative networks to underrepresented regions.
3.2 Machine Learning Algorithms
A total of 59 unique ML algorithms were identified across the 202
studies, with 475 algorithm mentions in total
(Fig. 4a; the
complete inventory, with the number of studies using each algorithm,
is given in Table S3). Artificial neural networks (ANN) were by far
the most widely adopted, appearing in 131 studies (64.9%). Random
forest (RF) was the second most popular
(
Regarding algorithm multiplicity, 45.5% of studies employed a
single algorithm, 53.5% used two or more (mean = 2.35, median = 2),
and the remaining two studies (1.0%) have no algorithm recorded at
all, an extraction gap of the kind quantified in
Section 2.6.2.
Multi-algorithm comparison studies typically benchmarked three to
five algorithms against each other to identify the best performer
for their specific adsorption system
(e.g. Refs. [11, 27, 28]). Regression was
the dominant task type (84.2% of studies), with 41.1% also
incorporating optimization components. Classification-based
approaches were relatively rare (
When evaluating reported best-performing models, ANN maintained
its leading position, being selected as the top model in 88 studies
(43.6%), followed by RF (
Temporal analysis of the top five algorithms revealed distinct adoption trajectories (Fig. 4c). ANN usage has remained consistently high since 2019, whereas ensemble methods have gained momentum since 2021. RF usage grew from one study in 2020 [32] to 18 in 2025, and XGBoost, virtually absent before 2022 [33, 34], rose to 15 studies in 2025 alone. This shift follows a broader trend in the ML community toward gradient boosting frameworks, which often achieve competitive performance with less hyperparameter tuning than deep neural networks. CatBoost and LightGBM have also appeared as alternatives in recent studies, while crystal graph convolutional neural networks (CGCNN) have so far been applied in a single study [35]. RSM usage has remained relatively stable, owing to its long-established role in experimental design optimization within chemical engineering.
Regarding software platforms, MATLAB was the most frequently
reported tool (31.2% of all studies), followed by Python (19.8%) and
Design-Expert (14.4%). A validation method could be recovered from
the text of 45.5% of studies, cross-validation being the most common
(
3.3 Emerging Contaminants
The reviewed studies addressed 12 distinct emerging contaminant
categories encompassing 146 unique compounds
(Fig. 5a; the 50
most frequently studied compounds are listed in Table S7).
Antibiotics were the most frequently studied category, appearing in
105 studies (52.0%), followed by pharmaceuticals
(
Beyond pharmaceuticals, endocrine-disrupting compounds (EDCs)
were targeted in 24 studies (11.9%), per- and polyfluoroalkyl
substances (PFAS) in 19 (9.4%), and microplastics in 12 (5.9%).
Pesticides (
Regarding co-occurrence, 47.6% of studies focused on a single EC category, while 52.4% addressed contaminants spanning two or more categories (mean = 1.8 categories per paper). This indicates a growing recognition that real-world water matrices contain mixtures of emerging pollutants, necessitating ML models capable of handling multi-contaminant scenarios.
At the individual compound level, tetracycline was the most
frequently modeled contaminant (
Temporal analysis of EC category trends
(Fig. 5c) revealed
that antibiotic adsorption modeling has grown consistently from 4
studies in 2019 to 30 in 2025, with pharmaceutical studies following
a similar trajectory. The most striking temporal trend was the
emergence of PFAS-related ML studies, absent before 2021 but rising
to 10 studies by 2025, which reflects heightened global regulatory
attention to “forever chemicals.” Microplastic studies appeared only
from 2023 onward (
3.4 Adsorbent Materials
Adsorbent information was available for 201 of the 202 studies
(99.5%), spanning 18 distinct material categories
(Fig. 6a). Biochar was
the most frequently studied adsorbent, appearing in 43 studies
(21.3%), reflecting its low cost, tunable surface chemistry, and
growing availability from agricultural and industrial waste streams.
Activated carbon ranked second (
Carbon-based materials collectively dominated the dataset. The four carbon categories account for 94 studies (46.5%) when each study is assigned to a single category, as in the counts above, and for 104 studies (51.5%) if studies combining a carbon material with another class are also counted; thirteen studies carry more than one adsorbent category and are the difference between the two figures. This dominance aligns with their high surface area, chemical modifiability, and established effectiveness for organic micropollutant removal. The remaining studies featured diverse materials spanning biosorbents, agricultural waste, polymers and resins, silica, and soil or sediment, indicating that ML modeling is being applied across the full spectrum of adsorbent technologies.
Temporal analysis revealed notable shifts in material preferences (Fig. 6b). Biochar usage grew sharply from 2022 onward, with 12 studies in 2024 and 13 in 2025, together 25 of the 43 biochar papers (58.1%). Activated carbon showed a similar recent surge, with 13 studies in 2025. The most striking trends were the emergence of MOF-based studies, which appeared only from 2023 [36] and reached 5 studies in 2025, and microplastic adsorption studies, which followed a parallel trajectory from 2023 through 2025. These trends mirror the broader materials science literature, where MOFs and microplastic–contaminant interactions are active research frontiers.
The cross-tabulation of adsorbent categories with EC categories (Fig. 6c) revealed clear material–contaminant pairing patterns. Biochar was predominantly applied to antibiotic adsorption (12 studies) and pharmaceutical removal (6 studies), consistent with its proven affinity for aromatic organic compounds. Activated carbon showed a similar antibiotic focus (7 studies) but was more evenly distributed across contaminant classes, including PFAS, EDCs, and heavy metals. Carbon nanotubes were applied exclusively to antibiotics and pharmaceuticals. PFAS removal studies showed no strong preference for a single adsorbent type, with papers distributed across biochar, activated carbon, MOFs, LDH, graphene oxide, polymer and resin materials, soil and sediment, and other sorbents, which reflects ongoing efforts to identify optimal sorbents for these recalcitrant compounds. EDC studies were concentrated in the biochar and graphene oxide categories, while pesticide adsorption was primarily modeled using activated carbon and biochar.
To provide a quantitative perspective on adsorption performance
across material types, the reported maximum adsorption capacity
(
3.5 Model Performance
Model performance was assessed primarily through the coefficient
of determination (
Among the 110 studies reporting an
Comparison of
Dataset size was reported in 78 studies (38.2%), revealing a
concerning distribution
(Fig. 8c). Nearly 40%
of those studies (
Small training sets are the most frequently voiced concern about
this literature, and the 202 studies allow it to be examined rather
than assumed. The relationship between
The correct inference from this is not reassurance. It is that
reported
3.5.1 Meta-analysis of algorithm performance
To move beyond descriptive comparison and provide a
statistically rigorous cross-study synthesis, we pooled the
reported
A funnel plot analysis was performed to assess potential
publication bias
(Fig. 10b). Visual
inspection revealed notable asymmetry, with a concentration of
studies reporting high
Subgroup analysis by adsorbent and contaminant categories
(Fig. 11) revealed
that model performance varies substantially across adsorption
systems. By contaminant category, pesticide and antibiotic studies
achieved the highest median
3.6 Feature Importance Synthesis Across Studies
While most ML models in the adsorption literature function as
“black boxes,” a growing subset of studies
(
Across the 15 SHAP-enabled studies, initial contaminant
concentration and BET surface area emerged as the two most
universally important features, each appearing as a top-ranked
predictor in all 15 studies with mean relative importance scores of
0.81 and 0.81, respectively
(Fig. 13a). This
finding is directly consistent with classical adsorption theory: the
Langmuir and Freundlich isotherms both predict strong dependence on
equilibrium concentration (
Beyond universal features, the synthesis revealed system-specific
predictors that carry mechanistic information
(Fig. 13). For
biochar-based adsorption, elemental ratios (H/C, O/C, O+N/C) and
pyrolysis temperature emerged as important predictors: features not
captured by classical isotherm models but directly related to
surface aromaticity and functional group density. For PFAS
adsorption, contaminant molecular weight and perfluoroalkyl chain
length were highly ranked [40, 27, 33],
reflecting the dominance of hydrophobic interactions and
size-dependent partitioning. For soil/sediment systems, organic
matter content and clay content outranked surface area
[49], consistent with the known role of soil organic
matter in contaminant sorption. When categorized by feature type,
contaminant properties (molecular weight, log
The strong alignment between ML-derived feature importance and classical adsorption theory validates the physical plausibility of these data-driven models. However, the synthesis also highlights that ML can extract insights beyond classical frameworks (particularly the importance of elemental ratios and molecular descriptors), suggesting a complementary role for explainable ML in hypothesis generation for adsorption mechanism research.
3.7 Machine-Extractable Reporting
To characterize how readily methodological detail can be
recovered from this literature at scale, each study was scored on
eight binary indicators (Table S5): whether a value could be
extracted for (1)
Recovery rates differed sharply between indicators
(Fig. 14a). Software
disclosure was the most readily extracted (61.9%), followed by
The score improved modestly over time
(Fig. 14b; Table S8;
Spearman
3.8 Explainable AI and the Distance to Deployment
The preceding sections establish which algorithms are used and
how well they are reported to perform. Neither question is the one a
utility engineer would ask, which is whether any of these models
could be relied upon outside the laboratory that produced it. Two
conditions bear on that: the model must be interpretable enough to
be trusted when it is wrong, and it must have been shown to work on
water resembling the water it would treat. The reviewed studies
speak to both, and in each case the answer is that the field is
earlier in its development than the reported
Interpretability is advancing but remains a minority practice. Fifteen studies (7.4%) apply SHAP, and their cross-study synthesis in Section 3.6 is encouraging: the features that emerge as most influential, initial contaminant concentration and BET surface area, are the ones classical adsorption theory would nominate, which is a meaningful, if modest, validation that these models are learning adsorption rather than the idiosyncrasies of a particular experiment. Its limits should be equally clear. A consensus drawn from 15 studies is a consensus among the small subset of authors who chose to look, and those authors were disproportionately working with tree ensembles, for which SHAP is computationally convenient. Whether the same features dominate in the 48 studies whose best model was a neural network is not established by these studies, and the agreement with theory does not by itself license using any of these models predictively.
Deployment is the larger gap. The water matrix was recorded for 132 studies (65%), and among these at least 46 (23% of all studies) report working with a real or non-synthetic matrix (wastewater, groundwater, drinking water, surface water, urine or tap water), while 86 (42%) report synthetic solutions alone. These counts should be read as lower bounds on real-matrix work, given the extraction audit’s finding that empty cells substantially reflect extraction failure rather than silence. Even read generously, the picture is one in which most models are trained and tested on single-contaminant solutions in deionised water, whereas the competitive adsorption, natural organic matter and variable ionic strength of a real influent are precisely what would degrade their predictions. That concern compounds with the validation practice documented in Section 3.5: a model fitted and evaluated on random splits of one synthetic dataset has not been tested against any of the conditions that deployment would impose on it.
Three further reporting deficits stand between this literature and reuse of its models by others. Software is identified in 125 studies (61.9%), with MATLAB the most common environment and Python-based stacks a growing minority, but the extracted dataset contains no field for code or data availability, so the number of studies releasing either cannot be stated here. Establishing it would require a separate full-text pass. Only 11 studies (5.4%) report both a dataset size and a number of input features, so the sample-to-feature ratio, the single most direct indicator of whether a model had enough data to support its complexity, is computable for one study in twenty. Among those 11 the median ratio is 47:1, with two studies below 10:1. And no reviewed study reports a prospective test, in which a model trained on one adsorbent–contaminant system predicts an experiment not yet performed. Until such tests become routine, the performance figures synthesised here should be read as descriptions of model fit, not as evidence of predictive capability in service.
3.9 Research Gaps and Future Directions
The synthesis of the 202 reviewed studies reveals several critical gaps that should guide future research in ML-assisted adsorption of emerging contaminants (Fig. 15).
3.9.1 Contaminant and material diversity
The current literature is heavily skewed toward pharmaceutical compounds, which account for 81.7% of all studies, while PFAS (9.4%), microplastics (5.9%), and pesticides (6.4%) remain substantially underrepresented. The top 10 individual contaminants appear in over 80% of studies, leaving hundreds of environmentally relevant emerging pollutants unstudied by ML approaches. On the adsorbent side, carbon-based materials (biochar, activated carbon, CNTs, graphene oxide) dominate at 46.5%, whereas high-performance materials such as MOFs (e.g. Refs. [36, 50, 40]), LDHs (e.g. Refs. [51, 44, 52]), and engineered nanocomposites (e.g. Refs. [42, 53, 54]) have received limited ML attention despite their growing prominence in adsorption research. Future studies should prioritize ML modeling for PFAS, microplastics, and multi-contaminant systems using a broader range of advanced adsorbents.
3.9.2 Data quantity and quality
Perhaps the most pressing concern is the prevalence of small datasets (Fig. 15b): 40% of studies that reported dataset size worked with fewer than 50 samples, and the median was only 131. Such limited training data raises questions about overfitting and generalizability, particularly for complex models like deep neural networks [47, 48]. Exceptions exist among studies that compiled large multi-source datasets (e.g. Refs. [11, 30, 35]), showing that aggregating published adsorption data across studies can yield training sets above 1,000 samples. Shared, open-access adsorption databases, analogous to those available in drug discovery or materials science, would accelerate that aggregation. A further 9.4% of studies reported the number of input features, which limits assessment of model complexity relative to dataset size.
3.9.3 Methodological rigor
Standardized reporting of performance metrics remains
inconsistent:
3.9.4 Algorithm innovation and frontier roadmap
While 59 unique algorithms were identified, ANN alone accounted for 64.9% of studies, although its share as the best-performing model has declined from 71% in 2019 to 32% in 2025. That comparison should be read with its base in view: 2019 contributed seven studies with a named best model, so 71% is five of seven, whereas the 2025 figure rests on 57 (Fig. 15a). More recent ML architectures, including graph neural networks (GNNs), transformers, and physics-informed neural networks (PINNs), were virtually absent from the reviewed literature, despite their demonstrated potential in related domains such as molecular property prediction and process modeling. Among deep learning approaches, only scattered examples of CNN [56, 57, 31], RNN [58], LSTM [59, 28], and DNN [47] were identified, while a crystal graph convolutional neural network (CGCNN) was applied in a single PFAS study [35]. The rapid adoption of XGBoost (from 0 studies before 2022 to 15 in 2025) and CatBoost (e.g. Refs. [11, 27, 40]) shows that the field is receptive to new methods, motivating a systematic assessment of emerging architectures.
Table 3 presents a
technology readiness assessment for six frontier ML paradigms with
high potential for adsorption modeling. Physics-informed neural
networks (PINNs) can embed thermodynamic constraints (e.g., Gibbs
free energy, mass balance) directly into the loss function,
ensuring physically plausible predictions even with limited data,
a critical advantage given that 40% of current studies use fewer
than 50 samples. Graph neural networks (GNNs) can represent
molecular structures and porous material topologies as graphs,
enabling structure–property predictions for novel adsorbents
without requiring hand-crafted descriptors; the sole CGCNN
application to PFAS sorption [35] demonstrated this
potential. Transfer learning offers a practical solution to the
small-dataset problem by pre-training models on large related
datasets (e.g., heavy metal adsorption,
| Technology | TRL | Impact | Adsorption Application | Key Challenge |
|---|---|---|---|---|
| Physics-Informed NN (PINN) | 2 | High | Embed isotherm/kinetic equations as constraints; enforce thermodynamic consistency | Formulating differentiable physics losses for heterogeneous systems |
| Graph Neural Networks (GNN) | 2 | High | Structure–property prediction for MOFs, polymers; molecular fingerprinting of ECs | Representing amorphous materials (biochar, AC); limited training data |
| Transfer Learning | 1 | High | Pre-train on heavy metal/dye datasets
( |
Domain shift between pollutant classes; negative transfer risk |
| Generative Models (VAE/Diffusion) | 1 | Medium | Inverse design of optimal adsorbent compositions for target contaminants | Synthesizability constraints; experimental validation loop |
| Bayesian ML | 3 | Medium | Uncertainty-aware predictions; active learning for experimental design | Computational cost; scalability to large feature spaces |
| Transformer / LLM | 1 | Medium | Multi-modal input (text + numeric); literature mining for automated data extraction | Hallucination risk; interpretability; data formatting |
3.9.5 Interpretability and mechanistic insight
Only 7.4% of studies employed model interpretability tools such as SHAP (SHapley Additive exPlanations), which leaves most ML models in this field functioning as “black boxes.” Studies employing ensemble methods with built-in feature importance analysis (e.g. Refs. [60, 61, 62]) provide partial interpretability, but systematic application of post-hoc explainability methods remains uncommon. Integrating explainability methods, including SHAP, LIME, partial dependence plots, and attention mechanisms, would increase trust in model predictions and generate mechanistic hypotheses about adsorption behavior. Coupling ML predictions with physicochemical understanding (e.g., surface functional groups, pore size distributions, electrostatic interactions) points toward physics-informed ML models, as exemplified by the QSAR/LFER-based approaches of Cho et al. [63, 37] and Lee et al. [38], which embed molecular descriptors into predictive frameworks.
3.9.6 Geographic and collaborative gaps
Research output is concentrated in a small number of countries, with China, India, Iran, and South Korea contributing 58.3% of all studies. Regions facing acute emerging contaminant challenges (including Sub-Saharan Africa [64, 65, 66], Southeast Asia [67], and Latin America [68, 69, 70, 71]) are underrepresented despite producing valuable individual studies. International collaborative networks and capacity-building initiatives could help democratize access to ML tools for water treatment research. Furthermore, the development of region-specific models trained on local water matrices and indigenous adsorbent materials would enhance practical applicability.
3.9.7 From prediction to application
The vast majority of reviewed studies focus on prediction accuracy as the primary objective, with limited attention to practical deployment. Studies that couple ML with response surface methodology for process optimization (e.g. Refs. [72, 73, 74]) or employ multi-objective optimization (e.g. Refs. [75, 76, 77]) take steps toward application-oriented modeling. Stacking and ensemble strategies designed for generalization [78, 79, 80] and Bayesian approaches that quantify prediction uncertainty [81, 82] offer pathways beyond point-estimate predictions. Future research should bridge this gap by developing ML-assisted decision-support tools for adsorbent selection (e.g. Refs. [83, 84]), real-time process optimization, and multi-objective optimization that balances removal efficiency, cost, and environmental impact. Integrating ML models with life cycle assessment (LCA) and techno-economic analysis would further support the translation of laboratory findings into scalable water treatment solutions.
3.10 Limitations of This Review
Five limitations bound what can be concluded from this synthesis, and one of them we regard as material.
3.10.1 Database coverage
The search drew on Scopus and PubMed only. Web of Science, IEEE Xplore and Engineering Village were not searched. The direction of the resulting bias can be reasoned about, though its size cannot be measured. Scopus and PubMed between them index the environmental-engineering and water-treatment journals in which adsorption work overwhelmingly appears, so studies motivated by a contaminant and an adsorbent are unlikely to have been missed systematically. The exposure is at the other end of the subject, among ML-motivated work published in computer-science and electrical-engineering venues, where IEEE Xplore has coverage the two databases used here do not. A study of that kind (an algorithmic contribution demonstrated on an adsorption dataset, published at an engineering venue) is precisely the kind this review would have failed to retrieve. No supplementary search was run to bound this, so the number of such studies is unknown. The literature reviewed here should be read as representative of the environmental-science work on ML for EC adsorption rather than of all literature in which such models appear. The consequence bears asymmetrically on the findings: the bibliometric and contaminant–adsorbent mapping would be little changed by the omission, whereas the algorithm-frequency picture, and in particular the dominance of ANN over more recent architectures, may understate methods favoured in computational venues.
3.10.2 Registration
The review was not prospectively registered and no protocol was published in advance, as set out in Section 2.2. The eligibility criteria and search query were fixed before the searches ran and were not amended, but this is a claim about our records rather than an independently verifiable commitment.
3.10.3 No study-level quality appraisal
The included studies were not appraised for risk of bias,
because no validated instrument exists for this study design and
the instruments built for clinical prediction models do not
transfer to it
(Section 2.7). The
consequence is that a study whose reported
3.10.4 Automated screening and extraction
Screening was performed by deterministic rule-based scripts and extraction by a regex pass reconciled against a language-model pass, with the accuracy of both measured by audit (Sections 2.6.1 and 2.6.2). The audits found no false exclusions in 101 sampled records and 93.3% accuracy in populated extraction cells, but they were conducted by the authors rather than by an independent assessor, they sample one screening stage rather than all three, and they establish that sparsely populated fields are unreliable as evidence of what studies do not report. Claims in this review that depend on such fields have been confined to the audited subset accordingly.
3.10.5 Search cut-off and the moving target
The searches ran on 24 February 2026 against a window closing at the end of 2025. Coverage of 2025 is therefore a lower bound, and work published in 2026 (including at least three reviews of adjacent or identical scope, as Table 1 records) falls outside the reviewed period by construction. This is inherent to any systematic review with a fixed cut-off, but in a field growing at the rate documented in Section 3.1 the interval between cut-off and publication carries more consequence than it would elsewhere.
4. Conclusions
This systematic review analyzed 202 peer-reviewed studies
(2015–2025) on machine learning for adsorption-based removal of
emerging contaminants from water, and is, to our knowledge, the first
review in this domain to combine a stated search and screening
protocol with a quantitative meta-analysis of reported performance and
a cross-study synthesis of feature importance. The field has grown
rapidly, with 179 of the 202 studies (88.6%) published from 2021
onward. Among 59 unique algorithms, ANN dominated (64.9%), though
ensemble methods are rapidly gaining adoption; inverse-variance
weighted meta-analysis revealed pooled
Cross-study synthesis of SHAP-based feature importance from 15 explainable ML studies identified initial concentration, BET surface area, and pH as universally important predictors, findings that align with classical Langmuir/Freundlich theory, while also revealing system-specific drivers (elemental ratios for biochar, molecular weight for PFAS) that extend beyond traditional frameworks. However, critical gaps persist: antibiotics, pharmaceuticals, NSAIDs and analgesics account for 81.7% of target contaminants while PFAS (9.4%) and microplastics (5.9%) remain underrepresented; 40% of the studies that report a dataset size used fewer than 50 data points; and validation, though nearly always performed, is nearly always an internal split of a single dataset rather than a test against independent data. The proposed MRCAS checklist (22 items across 5 categories) provides a practical tool to standardize future reporting.
To advance from prediction-focused modeling toward actionable decision-support tools, we recommend five priority directions: (1) diversifying contaminant coverage toward PFAS, microplastics, and multi-contaminant mixtures; (2) developing shared open-access adsorption databases to address the small-dataset bottleneck; (3) adopting the MRCAS reporting framework with rigorous external validation; (4) systematically evaluating frontier architectures (particularly physics-informed neural networks and graph neural networks) that can embed domain knowledge and handle molecular structures; and (5) expanding the use of explainable ML to bridge data-driven predictions with mechanistic understanding of adsorption processes.