Abstract:
Polymer science is undergoing a profound paradigm shift from labor-intensive, empirical trial-and-error experimentation to AI-guided, data-driven, and predictive material design. Owing to their multiscale structural diversity (spanning atomic connectivity, chain-packing behavior, and macroscopic morphology), polymer materials possess an enormous chemical design space that traditional experimental and computational methods cannot efficiently explore. High-precision quantum-chemical calculations are too computationally expensive for large-scale screening, whereas classical molecular dynamics simulations often involve trade-offs between force-field accuracy and accessible time scales. Against this backdrop, polymer informatics has emerged as a transformative interdisciplinary field that leverages artificial intelligence algorithms to mine structure–property–processing relationships from accumulated experimental, simulation, and literature data, enabling accurate property prediction and on-demand inverse design of polymer materials. This Review systematically summarizes recent advances in AI-guided informatics for polymer science based on an analysis of representative literature published between 2023 and 2025. The survey covers the current state of polymer informatics across four interconnected dimensions, organized around the logical backbone of data foundation, property prediction, inverse design, interpretability, and physical constraints. In the domain of polymer representation and data strategies, we evaluate the essential differences and applicable boundaries of major representation methods, including SMILES-based variants (PSMILES, BigSMILES, and HAPPY), molecular graph representations (including the hierarchical polymer graph, HPG), and topology-based multi-cover persistence (MCP) descriptors. We further discuss data augmentation techniques and transfer learning strategies developed to address the critical challenge of scarce high-quality labeled data and highlight the growing trend of multi-representation fusion for capturing both local chemical environments and global topological features. In the domain of machine learning for polymer property prediction, we review models and strategies for predicting thermodynamic properties (e.g., glass transition temperature (
Tg)), mechanical properties (tensile and flexural moduli), transport properties (gas permeability and diffusivity), electrical properties (dielectric constant, breakdown strength, and energy storage density), and functional service performance. In particular, we examine the complementary roles of graph neural networks (GNNs), Transformer-based architectures, large language models (e.g., PolyBERT, PolyNC), and multimodal learning frameworks in building accurate forward mappings from polymer structure to target properties and critically compare their performance-cost trade-offs. In the domain of generative models and inverse design, we discuss how variational autoencoders (VAEs), generative adversarial networks (GANs), reinforcement learning, and Bayesian optimization are integrated into closed-loop design frameworks that couple structure generation, property prediction, synthesizability evaluation, and experimental validation. We highlight emerging system-level approaches that move beyond single-objective generation toward multi-constraint, application-oriented polymer discovery, including de novo design of antimicrobial polymers, recyclable vitrimeric polymers, and high-performance dielectric materials. In the domain of model interpretability and physics-informed modeling, we review how attention mechanisms, SHAP (SHapley Additive exPlanations) analysis, and physics-informed neural networks are being applied to transform AI models from opaque black boxes into scientifically interpretable tools that can provide chemical insights and respect fundamental physical laws (e.g., thermodynamic constraints and scaling laws). Finally, we identify key challenges and outline five future research directions: (1) polymer data standardization and the construction of unified, community-shared benchmark datasets; (2) coupled physicochemical process modeling that integrates molecular-scale predictions with manufacturing-scale simulations; (3) multi-scale modeling bridging atomistic, mesoscale, and macroscopic descriptions; (4) generative artificial intelligence for creative, beyond-training-distribution polymer discovery; and (5) application-oriented material design with explicit consideration of synthesizability, processability, recyclability, and lifecycle performance. This Review aims to provide comprehensive references and valuable insights for researchers working at the intersection of polymer science, materials informatics, and artificial intelligence.