Thermo-Mechanical-Metallurgical Surrogates
A proof of concept Thermo-Mechanical-Metallurgical(TMM) machine learning model
using MALAMUTE/MOOSE and libtorch. The complete implementation will be in C++
for a fully automated pipeline for solving models and in-situ training of a
neural network with the results.
The initial models will be based on JMAK implementation from kmC-FEA.
Solver framework decision
Original plan: deal.ii + libtorch using
adamantine (ORNL) as the base AM
solver.
Revised plan: Use MALAMUTE/MOOSE (INL) as the base solver instead.
Why not adamantine
adamantine is a thermomechanical code for AM based on deal.II, ArborX, Trilinos, and Kokkos. After detailed analysis (see adamantine vs MALAMUTE):
- One-way coupling only: Temperature feeds stress, but deformation does not affect thermal field. No bidirectional TMM coupling.
- No metallurgical solver: Only post-processing of G, cooling rate, R at solidification front. No JMAK, KM, TRIP, or CCT integration.
- Hard to extend: Heavy template parameterization (
dim,n_materials,p_order,MaterialStates,MemorySpaceType) makes adding new physics modules difficult. No plugin system or Python bindings. - No explicit solver for coupled systems: Explicit time stepping only (implicit removed in v1.1). Fully coupled TMM requires implicit.
- Mechanical MPI without AMR: v1.1 added MPI for mechanical but AMR + MPI for mechanical is still ongoing.
Why MALAMUTE/MOOSE
MALAMUTE (MOOSE Application Library for Advanced Manufacturing UTilitiEs) built on the MOOSE framework from Idaho National Lab:
- Fully coupled, fully implicit multiphysics: All physics solved in one Newton system. Bidirectional TMM coupling is native.
- Easy custom physics:
ADKernel/ADMaterialwith automatic Jacobians via MetaPhysicL. ~100 LOC per kernel. JMAK, KM, TRIP can be implemented as custom MOOSE objects. - Built-in ML infrastructure:
stochastic_toolsmodule provides GP surrogates, active learning, LibTorch integration (load TorchScript models for inference), POD ROM, and UQ sampling. - MultiApp multiscale: Nested sub-applications enable running microstructure models at quadrature points while macro solve runs on parent app.
- PETSc solver infrastructure: Direct (SuperLU_dist), iterative (full KSP suite), Newton, JFNK, adaptive time stepping.
- Melt pool CFD: INS with level set, Marangoni, recoil pressure (adamantine has none of this).
- No recompilation for experiments: HIT input files allow rapid parameter sweeps.
What is lost
- Kokkos GPU portability (adamantine has better GPU path)
- Built-in EnKF data assimilation from experimental thermography
- deal.II’s hp-adaptivity (h/p refinement vs h-only in MOOSE)
Mitigation
- GPU: PETSc is adding GPU backends (Kokkos, CUDA, HIP)
- EnKF: Can implement within MOOSE’s stochastic_tools or use external library
- hp-adaptivity: h-refinement sufficient for AM thermal problems
Implementation plan with MALAMUTE
- Phase 1: Implement JMAK/KM as MOOSE
ADKernelobjects for diffusional and martensitic transformations - Phase 2: Implement Leblond TRIP as
ADMaterialwith volume change eigenstrain - Phase 3: Train GNO offline on MALAMUTE FEM solutions (Python)
- Phase 4: Export GNO as TorchScript, deploy via MOOSE’s LibTorch surrogate infrastructure
- Phase 5: Implement in-situ training loop with uncertainty- driven retraining
Related research notes
- MOOSE
- Non-Isothermal JMAK Phase Transformation Kinetics
- TRIP and Transformation Plasticity in Welding
- Microstructure Evolution Modeling for AM
- Neural Operators for PDE Surrogates
- PINNs for Transient Thermal Problems
- Reduced-Order Modeling with ML Surrogates
- In-Situ Training and Active Learning for Surrogates
- deal.II and LibTorch Integration Patterns
Topics and avenues to explore
Metallurgical modeling (new domain)
Phase transformation kinetics:
- JMAK theory extensions beyond isothermal conditions
- Koistinen-Marburger model for martensitic transformations
- Continuous cooling transformation (CCT) diagram integration
- Multi-phase field approaches vs. empirical kinetics
- Keywords:
non-isothermal JMAK kinetics,CCT diagram FEM integration,Koistinen-Marburger welding,phase field additive manufacturing
Microstructure evolution:
- Grain growth models (Monte Carlo Potts, cellular automata)
- Texture evolution during WAAM thermal cycles
- Precipitation kinetics (Kampmann-Wagner numerical model)
- CALPHAD-coupled simulations (Thermo-Calc, Pandat integration)
- Keywords:
Monte Carlo Potts grain growth welding,cellular automata microstructure AM,Kampmann-Wagner precipitation,CALPHAD FEM coupling
Transformation-induced effects:
- Transformation-induced plasticity (TRIP) models
- Volume change from phase transformations and stress coupling
- Greenwood-Johnson mechanism for transformation plasticity
- Keywords:
TRIP model welding FEM,Greenwood-Johnson transformation plasticity,phase transformation volume change stress
ML surrogate modeling (new domain)
Neural operators for PDE surrogates:
- Fourier Neural Operators (FNO) for spatio-temporal fields
- DeepONet architecture for operator learning
- Graph Neural Operators for unstructured meshes
- Low-rank factorization for high-dimensional outputs
- Keywords:
Fourier neural operator PDE,DeepONet surrogate model,graph neural operator FEM mesh,neural operator heat equation
Physics-informed approaches:
- PINNs for transient heat conduction with moving sources
- Conservative neural networks (energy/mass conservation)
- Hard constraint enforcement vs. penalty methods
- Hybrid FEM-ML: ML for closure terms, FEM for conservation
- Keywords:
PINN transient heat conduction,conservation neural network,hard constraint PINN,hybrid FEM machine learning
Reduced-order modeling:
- Proper Orthogonal Decomposition (POD) for thermal fields
- Autoencoder-based dimensionality reduction
- POD-Galerkin projection with neural network closure
- Dynamic mode decomposition (DMD) for transient behavior
- Keywords:
POD thermal field reduction,autoencoder FEM surrogate,POD neural network closure,dynamic mode decomposition transient
In-situ training and active learning (new domain)
Online training during simulation:
- Training surrogate concurrently with FEM solver
- When to trigger training (error threshold, iteration count)
- Memory-efficient data pipelines for long simulations
- Checkpointing and resuming training state
- Keywords:
online surrogate training simulation,concurrent FEM machine learning,adaptive surrogate model training
Active learning for simulation data:
- Uncertainty-driven sampling (where surrogate is uncertain)
- Query-by-committee, expected improvement strategies
- Adaptive mesh refinement guided by surrogate error
- Multi-fidelity: coarse FEM for exploration, fine for refinement
- Keywords:
active learning FEM sampling,uncertainty-driven surrogate training,multi-fidelity surrogate additive manufacturing,Bayesian optimization simulation
Error estimation and UQ:
- Ensemble methods for epistemic uncertainty
- Conformal prediction for surrogate confidence bounds
- A posteriori error estimation for neural operators
- Propagation of surrogate error to downstream predictions
- Keywords:
ensemble uncertainty neural operator,conformal prediction surrogate,a posteriori error neural network,surrogate error propagation
MALAMUTE/MOOSE + libtorch integration (revised from deal.ii)
Note: Originally planned with deal.ii + libtorch. Revised to MALAMUTE/MOOSE after analysis. See Adamantine vs MALAMUTE and deal.II + LibTorch Integration for detailed comparison.
MOOSE patterns for coupled problems:
- ADKernel/ADMaterial with automatic Jacobians for custom physics
- MultiApp system for nested sub-applications at different scales
- PETSc solver infrastructure (SuperLU_dist, KSP, GAMG)
- MPI domain decomposition via libMesh (scales to >100k cores)
- Keywords:
MOOSE ADKernel custom physics,MOOSE MultiApp multiscale,PETSc block preconditioner AM
libtorch embedded in MOOSE simulation:
- MOOSE’s
stochastic_toolsmodule has native LibTorch integration - Load TorchScript models for inference as surrogate materials/BCs
- Model loading and inference within solver loops
- Gradient computation for in-situ training
- Memory management for large tensor operations
- Keywords:
MOOSE LibTorch surrogate,libtorch C++ embedded inference,TorchScript MOOSE integration
Literature review plan
Phase 1: Metallurgical foundations (weeks 1-2)
Goal: Build domain knowledge in phase transformations and microstructure modeling
-
JMAK and phase transformation kinetics:
- Search:
("JMAK" OR "Johnson-Mehl-Avrami-Kolmogorov") AND ("non-isothermal" OR "continuous cooling") AND ("welding" OR "additive manufacturing") - Focus: How JMAK is adapted for non-isothermal thermal cycles, limitations, extensions
- Output: Summary of JMAK variants and their applicability to WAAM thermal histories
- Search:
-
CALPHAD-integrated simulations:
- Search:
("CALPHAD" OR "Thermo-Calc") AND ("finite element" OR FEM) AND ("welding" OR "additive manufacturing") - Focus: How thermodynamic databases couple to FEM, computational cost, accuracy gains
- Output: Map of CALPHAD-FEM coupling strategies and software implementations
- Search:
-
Transformation-induced plasticity:
- Search:
("TRIP" OR "transformation induced plasticity") AND ("welding" OR "additive manufacturing") AND ("finite element" OR FEM) - Focus: When TRIP matters (material systems, thermal cycles), model formulations
- Output: Decision tree for when to include TRIP in TMM models
- Search:
Phase 2: ML surrogates for physics simulation (weeks 3-4)
Goal: Understand the landscape of neural operators and physics-informed ML
-
Neural operator architectures:
- Search:
("Fourier neural operator" OR DeepONet) AND ("PDE" OR "partial differential equation") AND ("heat transfer" OR "solid mechanics") - Focus: Architecture choices for spatio-temporal problems, generalization across geometries
- Output: Comparison table of neural operator architectures for thermal/mechanical PDEs
- Search:
-
Physics-informed neural networks:
- Search:
("physics-informed neural network" OR PINN) AND ("transient" OR "time-dependent") AND ("heat equation" OR "heat conduction") - Focus: Training stability for transient problems, handling moving sources, scalability
- Output: PINN limitations and success factors for transient thermal problems
- Search:
-
Reduced-order modeling with ML:
- Search:
("proper orthogonal decomposition" OR POD) AND ("neural network" OR "machine learning") AND ("finite element" OR FEM) - Focus: POD-NN hybrids, when ROM is sufficient vs. when neural operators needed
- Output: Decision framework for choosing ROM vs. neural operator approach
- Search:
Phase 3: In-situ training and active learning (week 5)
Goal: Understand how to train surrogates during simulation, not post-hoc
-
Online/concurrent training:
- Search:
("online training" OR "in-situ training" OR "on-the-fly") AND ("surrogate model" OR "reduced order") AND ("simulation" OR "FEM") - Focus: Training triggers, data selection strategies, convergence criteria
- Output: Taxonomy of in-situ training approaches and their computational overhead
- Search:
-
Active learning for simulation:
- Search:
("active learning" OR "adaptive sampling") AND ("surrogate" OR "emulator") AND ("computational model" OR "simulation") - Focus: Query strategies, uncertainty metrics, multi-fidelity approaches
- Output: Active learning strategy recommendations for FEM surrogate training
- Search:
-
Error estimation and UQ:
- Search:
("uncertainty quantification" OR UQ) AND ("neural operator" OR "surrogate model") AND ("PDE" OR "partial differential equation") - Focus: How to quantify surrogate confidence, error propagation to predictions
- Output: UQ methods suitable for TMM surrogate predictions
- Search:
Phase 4: MALAMUTE/MOOSE + libtorch implementation patterns (week 6)
Goal: Understand how to integrate ML into MOOSE-based simulations (revised from deal.ii)
-
MOOSE for coupled thermal-mechanical-metallurgical problems:
- Search:
MOOSE ("thermal" OR "heat transfer") AND ("additive manufacturing" OR "welding") - Focus: MALAMUTE melt pool modules, MultiApp multiscale patterns, ADKernel for custom physics
- Output: MOOSE modules and examples most relevant to TMM implementation
- Search:
-
libtorch in MOOSE via stochastic_tools:
- Search:
MOOSE ("libtorch" OR "PyTorch C++" OR "TorchScript") AND ("surrogate" OR "neural network") - Focus: MOOSE’s native LibTorch integration, surrogate material models, performance overhead
- Output: Reference implementations of LibTorch surrogates in MOOSE simulations
- Search:
-
Hybrid FEM-ML architectures:
- Search:
("hybrid" OR "coupled") AND ("machine learning" OR "neural network") AND ("finite element" OR FEM) AND ("closure" OR "surrogate") - Focus: Where to insert ML (constitutive laws, boundary conditions, full field prediction)
- Output: Architecture options for TMM surrogate with clear tradeoffs
- Search:
Phase 5: Gap analysis (week 7)
Goal: Synthesize findings into specific research gap identification
-
TMM + ML intersection matrix:
- Rows: thermal-only, thermo-mechanical, thermo-metallurgical, full TMM
- Columns: analytical, FEM, ROM, neural operator, hybrid FEM-ML
- Mark existing work, identify empty cells as gaps
-
In-situ training gap:
- How many works train surrogates during simulation vs. post-hoc?
- What problems have been solved with in-situ training?
- What is missing for TMM problems specifically?
-
Computational efficiency gap:
- Tabulate: model fidelity vs. computational cost for existing TMM approaches
- Where can surrogates provide 10x-100x speedup without accuracy loss?
- What is the bottleneck in current TMM simulations?
-
Open-source implementation gap:
- How many TMM + ML works are open-source?
- What frameworks are used (commercial vs. open-source)?
- Where is there opportunity for a MALAMUTE/MOOSE + libtorch reference implementation?
Future exploration: complete fast TMM prediction model
Architecture options for chained prediction
Option 1: GNN → PINN chain
GNN predicts temperature and strain histories from tool path, feeds into PINN for metallurgical prediction. Feasible but error propagation is the critical flaw. Metallurgical kinetics are non-linear in temperature and strain, so upstream errors compound non-linearly through the JMAK/KM equations. The PINN enforces physics on metallurgical equations but cannot correct approximate inputs.
Requires autoregressive or recurrent GNN (GraphGRU, EvolveGCN) to capture history dependence, not static GNN.
Option 2: Single PINN for all three physics
Low feasibility. Three coupled PDE systems in one loss function: heat equation (parabolic, moving source), momentum balance (elliptic with plasticity), kinetics ODEs (stiff, path-dependent). Training instability from different equation scales, stiffness, and convergence rates. PINNs already struggle with multi-term loss balancing for single PDEs.
Path dependence requires either augmenting input space with time (4D spatio-temporal) or recurrent PINN variants, both unproven at this complexity.
Option 3: Neural operator chain (recommended)
Three-stage architecture:
-
Graph Neural Operator for thermal field:
- Input: tool path + process parameters
- Output: full temperature history T(x,t)
- Trained on FEM thermal solutions
-
Graph Neural Operator for mechanical field:
- Input: T(x,t) from stage 1 + tool path
- Output: strain/stress history ε(x,t), σ(x,t)
- Can be trained jointly with stage 1 to reduce error propagation
-
Local kinetics solver (not PINN):
- Input: T(x,t) and ε(x,t) histories at each material point
- JMAK/KM equations are ODEs at each point, decoupled spatially once T and ε are known
- Solve analytically or with small MLP approximating ODE integration
- Computationally cheap, no PINN needed
Why neural operator chain beats PINNs
- Neural operators learn the mapping directly, no loss balancing across physics
- Path dependence handled naturally by learning full spatio-temporal operator
- Metallurgical step is local (ODE per material point), no PDE-constrained network needed
- Training is sequential and stable
Error propagation mitigation
- Joint training: Train thermal and mechanical operators with combined loss including downstream metallurgical error, not just field-level MSE
- Uncertainty quantification: Use ensemble or Bayesian GNO to propagate uncertainty bounds through the chain
- Multi-fidelity correction: Train small correction network on high-fidelity FEM data that adjusts chain output
Literature status (updated June 2026)
- Neural operators for thermal fields in AM: demonstrated (FNO for LPBF thermal by Liu et al. 2024, GNO for WAAM by Chen et al. 2026)
- Neural operators for thermo-mechanical: emerging (PIDeepONet-RNN by Tian et al. 2025 for AM distortion)
- GNO-DeepONet hybrid for coupled TMM: Chen et al. (2026) in Additive Manufacturing — closest existing work to this project
- PINNs for coupled TMM: no successful demonstrations for full coupling; works only for simplified single-physics cases (see PINNs for Transient Thermal Problems)
- In-situ training of chained operators: unexplored, this is the gap (see In-Situ Training and Active Learning)
- No open-source TMM + ML reference implementation exists (gap for MALAMUTE + LibTorch)