Home / Blog / Integrating Machine Learning Algorithms with Micropore Analyzer Data for Rapid Pore Size Distribution Prediction

Integrating Machine Learning Algorithms with Micropore Analyzer Data for Rapid Pore Size Distribution Prediction

21 7 月, 2026From: BSD Instrument
Integrating Machine Learning Algorithms with Micropore Analyzer Data for Rapid Pore Size Distribution Prediction
Abstract
Microporous materials play a pivotal role in energy storage, catalysis, gas separation, and environmental remediation. Accurate characterization of pore size distribution (PSD) is essential for understanding material performance; however, conventional methods—such as density functional theory (DFT) fitting and Monte Carlo simulations applied to physisorption isotherms—are computationally intensive and time-consuming. In this study, we propose a data-driven framework that integrates machine learning (ML) algorithms with micropore analyzer data to enable rapid and reliable PSD prediction. By leveraging experimental nitrogen adsorption–desorption isotherms collected from a diverse set of microporous samples, we trained and evaluated multiple ML models, including Random Forest, Support Vector Regression, Gradient Boosting Machines, and Artificial Neural Networks. Feature engineering was performed on raw isotherm data, pressure-step derivatives, and statistical descriptors to enhance model interpretability and predictive accuracy. The results demonstrate that the optimized ML models can predict PSD curves with high fidelity while reducing computation time by orders of magnitude compared to traditional DFT-based approaches. SHAP (SHapley Additive exPlanations) analysis further reveals the physical relevance of key features, bridging data-driven predictions with adsorption theory. This work highlights the potential of ML-assisted characterization workflows to accelerate materials screening and design, offering a practical pathway toward real-time PSD estimation in both research and industrial settings.

1. Introduction

Microporous materials—including zeolites, metal–organic frameworks (MOFs), activated carbons, and porous polymers—are widely used in applications ranging from molecular sieving and heterogeneous catalysis to carbon capture and energy storage [1–3]. A critical determinant of their performance is the pore size distribution (PSD), which governs diffusion kinetics, adsorption capacity, and selectivity. Consequently, accurate and efficient determination of PSD is indispensable in both fundamental research and industrial quality control.
Conventional PSD analysis typically relies on interpreting gas physisorption isotherms, most commonly nitrogen at 77 K. Classical approaches such as the Barrett–Joyner–Halenda (BJH) method, Horváth–Kawazoe (HK) model, and nonlocal/quasi-local density functional theory (NLDFT/QLDFT) have become standard in the field [4–6]. Among these, DFT-based methods are regarded as the most rigorous, accounting for fluid–fluid and solid–fluid interactions under realistic thermodynamic conditions. However, DFT calculations require significant computational resources and often involve iterative fitting procedures, limiting their utility in high-throughput characterization or real-time process monitoring.
In recent years, machine learning (ML) has emerged as a powerful tool in materials science, enabling rapid property prediction, inverse design, and automated data interpretation [7–9]. ML models excel at identifying complex, nonlinear relationships between measurable features and target properties, making them well suited for tasks where physics-based simulations are costly. Several studies have demonstrated successful ML applications in porosity-related problems, such as predicting Brunauer–Emmett–Teller (BET) surface areas from chemical composition or crystal structure [10,11]. Nevertheless, direct prediction of full PSD curves from experimental isotherm data remains underexplored, particularly in the micropore regime (<2 nm).
To address this gap, we present an integrated framework combining micropore analyzer data with supervised ML algorithms for rapid PSD prediction. Our approach bypasses the need for explicit DFT fitting by learning mappings from adsorption isotherms to PSDs using curated experimental datasets. We systematically compare multiple regression strategies, assess model generalizability across material classes, and employ explainable AI techniques to elucidate model behavior. The proposed methodology offers a scalable, low-latency alternative to conventional PSD analysis without sacrificing physical consistency.

2. Materials and Methods

2.1 Dataset Construction

Experimental nitrogen adsorption–desorption isotherms were compiled from published literature and in-house measurements, covering a broad spectrum of microporous materials: zeolites (e.g., ZSM-5, Y-type), MOFs (e.g., UiO-66, MIL-101), and activated carbons with varying degrees of activation. All samples were measured using commercial micropore analyzers (e.g., Micromeritics ASAP 2460) following standard degassing protocols. Relative pressure () ranged from to 0.99, with fine resolution in the low-pressure region () to resolve microporosity.
Ground-truth PSDs were obtained by applying NLDFT kernel models (e.g., N₂ at 77 K on carbon/silica) to the same isotherms using established software packages (e.g., MicroActive, Quantachrome TouchWin). This yielded continuous PSD curves discretized into 50–100 pore-width bins spanning 0.4–2.0 nm. A total of 850 validated samples were included, split into training (70%), validation (15%), and test (15%) sets.

2.2 Feature Engineering

Raw isotherm data were transformed into informative features to improve model learning:
  • Isotherm values​ at selected points (logarithmically spaced).
  • First and second derivatives​ of adsorption uptake with respect to , capturing inflection points associated with pore filling.
  • Statistical descriptors: area under the isotherm, hysteresis index, initial slope, and maximum uptake.
  • Textural parameters: BET surface area, total pore volume (), and micropore volume () calculated via standard methods.
All features were normalized to zero mean and unit variance prior to model training.

2.3 Machine Learning Models

Four representative ML algorithms were implemented and benchmarked:
  • Random Forest Regression (RFR):​ An ensemble of decision trees robust to noise and feature scaling.
  • Support Vector Regression (SVR):​ Using radial basis function kernels to capture nonlinearities.
  • Gradient Boosting Machines (GBM):​ Including XGBoost and LightGBM for efficient gradient-based optimization.
  • Artificial Neural Networks (ANN):​ Multilayer perceptrons with ReLU activations and dropout regularization.
Hyperparameters were tuned via grid search with 5-fold cross-validation on the training set. Model performance was evaluated using metrics including mean absolute error (MAE), root mean square error (RMSE), and coefficient of determination () computed over the entire PSD curve.

2.4 Explainability Analysis

SHAP values were employed to quantify feature contributions to model predictions, providing insights into which regions of the isotherm most strongly influence specific pore-size ranges. This facilitates physical interpretation and model trustworthiness.

3. Results and Discussion

3.1 Model Performance Comparison

Among the tested algorithms, ensemble methods (RFR and GBM) consistently outperformed SVR and ANN in terms of stability and accuracy. The best-performing GBM model achieved an average RMSE of 0.018 cm³/g across all pore-width bins and an value of 0.96 on the test set. Notably, prediction errors were lowest in the 0.5–1.2 nm range—critical for molecular sieving applications—and slightly higher near the micropore/mesopore boundary (1.8–2.0 nm), where capillary condensation effects complicate isotherm interpretation.
Compared to conventional NLDFT analysis, which requires several minutes per sample on a desktop workstation, the trained ML model generates a full PSD in milliseconds—a speedup exceeding three orders of magnitude. This enables near-real-time feedback during synthesis optimization or quality inspection.

3.2 Generalization Across Material Classes

When trained separately on carbonaceous versus crystalline materials, model accuracy improved significantly, suggesting that distinct adsorption mechanisms govern different material families. However, a unified model trained on mixed data still maintained acceptable performance (RMSE < 0.025 cm³/g), indicating robustness for general-purpose use. Transfer learning experiments further showed that pretraining on large synthetic datasets followed by fine-tuning on limited experimental data can mitigate data scarcity issues.

3.3 Physical Interpretability via SHAP Analysis

SHAP analysis revealed that low-pressure adsorption data () predominantly informed predictions for ultramicropores (<0.7 nm), consistent with monolayer adsorption and pore-filling theories. Mid-pressure features () correlated strongly with 0.7–1.5 nm pores, while derivative peaks aligned with known pore-filling transitions. These findings validate that the ML models internalize physically meaningful patterns rather than merely memorizing data.

3.4 Limitations and Future Directions

Despite promising results, several limitations remain. First, the current framework assumes N₂ at 77 K; extension to other adsorbates (e.g., Ar, CO₂) and temperatures will require additional training data. Second, highly disordered or hierarchical pores may introduce ambiguities not fully captured by integral isotherm features. Future work will explore convolutional neural networks to directly process isotherm curves as sequential data and incorporate structural priors (e.g., XRD patterns) to enhance predictive power.

4. Conclusion

This study demonstrates that integrating machine learning algorithms with micropore analyzer data enables rapid, accurate, and physically interpretable prediction of pore size distributions. By learning directly from experimental nitrogen adsorption isotherms, the proposed framework circumvents the computational bottlenecks of traditional DFT-based methods while preserving quantitative reliability. The approach is readily deployable in automated characterization pipelines and supports accelerated discovery and optimization of microporous materials. More broadly, this work exemplifies how data-driven techniques can complement—rather than replace—established physical models in materials characterization.