Interdisciplinary AI: starting with questions from other fields

AI can help chemists screen molecules and biologists study cells. It is also used in weather forecasting, mathematical proofs and financial modelling. Researchers are exploring how AI can choose training configurations, run experiments and improve code.

How AI contributes to research

Read about this field ↓

Chemistry and materials: Predict properties, plan reactions, and find new materials

Will a molecule dissolve readily in water? Is a crystal stable? What will a set of reactants produce? AI can predict properties and screen candidates from structures and experimental records. It can also learn energies and forces to reduce the computational cost of molecular dynamics.

Property prediction

Represent a molecule as a graph of atoms and bonds, or as a structure with three-dimensional coordinates, to predict quantities such as solubility and energy. MPNN is a classic introduction to graph learning for quantum chemistry.

Reactions and synthesis

Forward prediction infers products from reactants; retrosynthesis works backward from a target molecule to find precursors and synthesis routes. Molecular Transformer models the SMILES strings of reactants and products as sequences.

Materials discovery and simulation

GNoME combines graph networks with first-principles calculations to screen stable crystals. Deep Potential learns potential energy to provide the energies and forces needed for molecular dynamics. Active learning uses existing results to select the next calculations or experiments.

  1. 2017MPNN

    Structure → properties

  2. 2018–2019Deep Potential / Molecular Transformer

    Simulation and reaction prediction

  3. 2023GNoME

    Model screening + DFT validation

2017 · ICML

Neural Message Passing for Quantum Chemistry

Neural Message Passing for Quantum Chemistry · Gilmer et al. · Fig. 1
Figure excerpt from the paper · Gilmer et al. · Fig. 1 · Original paper

The problem

How can atoms and their relationships predict molecular properties?

The key idea

Update node representations through messages, then aggregate them for a molecular prediction.

Why this paper

It grounds graph learning in a concrete chemistry task.

Where to start

Start with message passing, readout, targets, and data.

2018 · Physical Review Letters

Deep Potential Molecular Dynamics: A Scalable Model with the Accuracy of Quantum Mechanics

Deep Potential Molecular Dynamics: A Scalable Model with the Accuracy of Quantum Mechanics · Zhang et al., Deep Potential Molecular Dynamics, Figure 1
Figure excerpt from the paper · Zhang et al., Deep Potential Molecular Dynamics, Figure 1 · Original paper

The problem

How can accurate energy and force calculations become cheaper?

The key idea

Learn local atomic contributions from first-principles data while preserving problem symmetries.

Why this paper

An example of replacing an expensive step in a scientific computation.

Where to start

Trace energies and forces into simulation, then inspect physical comparisons.

A question to keep asking

Accuracy depends on covered configurations and reference calculations; new conditions need testing.

Papers and learning resources

SMILES
A string representation of molecular structure.
DFT
Density functional theory: a family of methods for calculating electronic structure, energy, and other properties.

Try it yourself

Use a solubility dataset from DeepChem to compare molecular fingerprints with traditional regression against a graph network. Try both random splits and molecular scaffold splits, and examine how the error changes on new scaffolds.

Physics and weather: Learn solutions to equations and predict how systems evolve

How fluids move, how heat spreads, and how the weather changes over the next few days usually involve physical quantities that vary in space and time. AI can approximate the process of solving these equations, or use observations to predict how physical systems evolve.

Bringing equations into training

Physics-informed neural networks (PINNs) train with equation residuals, boundary conditions, and observation errors together. Research examines how these constraints help, and how models handle complex domains and changes across multiple scales.

Learning solutions for a family of problems

Neural operators learn mappings from functions to functions, such as predicting an entire flow field from initial conditions. FNO constructs operations in the frequency domain, trains on solutions to a set of problems, and then handles new input conditions.

Forecasting weather and uncertainty

GraphCast uses graph networks to predict global weather fields. Later probabilistic forecasting approaches describe uncertainty through multiple possible futures. Researchers compare forecast errors, performance on extreme weather, and computation time.

  1. 2019PINNs

    Equation constraints in training

  2. 2021FNO

    Learning mappings between functions

  3. 2023GraphCast

    Predicting global weather fields

2021 · ICLR

Fourier Neural Operator for Parametric Partial Differential Equations

Fourier Neural Operator for Parametric Partial Differential Equations · Figure 1
Figure excerpt from the paper · Figure 1 · Original paper

The problem

How can a set of initial conditions directly predict continuous physical fields?

The key idea

Parameterize integral kernels in the frequency domain to learn mappings from functions to functions for a family of partial differential equation problems.

Where to start

Start with the operator layer and the inputs and outputs, then compare errors across resolutions and equation settings.

Papers and learning resources

PDE
A partial differential equation describes how a quantity changes with variables such as space and time.
Neural operator
A model that learns a mapping from an input function to an output function, such as an initial field to a future flow field.

Try it yourself

Start with a one-dimensional diffusion or Burgers equation. Use numerical solutions as a reference, compare model predictions under different initial conditions, and plot how the error changes over time.

Proteins and enzymes: From sequence to structure to functional design

Proteins consist of amino acid sequences, and their folded structures are closely linked to what they can do. AI can predict structures, find natural sequences with desired properties, or propose new designs.

Representation and structure prediction

Protein language models learn transferable features from sequences. AlphaFold2 combines sequences, evolutionary information, and geometric relationships to predict three-dimensional structures. AlphaFold3 extends this to complexes containing proteins, nucleic acids, small molecules, and other components.

Sequence and backbone design

ProteinMPNN designs sequences for a given backbone; RFdiffusion generates backbones that satisfy constraints; ESM3 jointly uses sequence, structure, and function information. The design workflow then checks expression, stability, binding, or catalytic activity.

Enzyme mining and engineering

Enzyme mining finds candidates in natural sequence libraries; enzyme engineering modifies existing enzymes; de novo design builds new candidates around a target reaction. VenusMine retrieves candidates using structural and sequence clues, while VenusRXN matches enzymes starting from chemical reactions.

  1. 2021AlphaFold2

    Structure prediction

  2. 2022–2023ProteinMPNN / RFdiffusion

    Sequence and backbone design

  3. 2025–2026AMix-1 / AMix-2

    Conditional generation and protein–text modeling

Read further

How enzyme engineering forms an experimental loop

Zero-shot scoring can first rank mutations. A small amount of measured data can fit a specific property, while active learning uses existing results and uncertainty to choose the next experiments. Each round adds newly measured activity or stability to the data and updates the candidate ranking.

Constraints in de novo design

One approach finds a backbone to support a catalytic core; another organizes local interactions within an existing backbone. Catalytic geometry, foldability, stability, and reaction conditions all determine which experiments are needed next.

New approaches to foundation models

AMix-1 models proteins with Bayesian flow networks and explores conditioning on multiple sequence alignments and inference-time search. AMix-2 unifies protein and text modeling, using causal generation between blocks and diffusion within each block. When reading, trace the roles of conditioning inputs, candidate generation, and verifiers separately.

2021 · Nature

Highly accurate protein structure prediction with AlphaFold

Highly accurate protein structure prediction with AlphaFold · Jumper et al., AlphaFold, Figure 1(e) · CC BY 4.0
Figure excerpt from the paper · Jumper et al., AlphaFold, Figure 1(e) · CC BY 4.0 · Original paper

The problem

How can sequence information support three-dimensional structure prediction?

The key idea

Combine evolutionary sequence information and residue relationships, iteratively refining representations and structure.

Why this paper

It connects data, geometry, and system design around a disciplinary problem.

Where to start

Start with the architecture, CASP14 evaluation, and confidence estimates.

2025 · arXiv

AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model

AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model · Figure 1
Figure excerpt from the paper · Figure 1 · Original paper

The problem

How can family information and verification feedback support protein design?

The key idea

Model proteins with Bayesian flow networks, use MSA context, and explore design workflows with larger inference-time search budgets.

Where to start

Trace generation conditions, verifiers, and experimental measurements separately, comparing their roles in the iterative process.

Papers and learning resources

MSA
Multiple sequence alignment aligns related protein sequences to reveal evolutionary conservation and variation.
Inverse folding
Finding an amino acid sequence that can adopt a given protein backbone structure.

Try it yourself

Choose a protein dataset with measured properties, and compare simple sequence features with frozen pretrained representations. Split the data by homology to test predictions on new families. For structure tasks, start by examining one AlphaFold prediction and its confidence plots.

Genomics: How sequence changes affect gene regulation

DNA variants can alter protein coding, but they can also affect regulatory elements, expression, or splicing. Genomic models try to turn these relationships within long sequences into predictable signals, helping researchers analyze what a variant may change.

From sequence to signals

Input a DNA sequence to predict signals associated with experiments, such as expression, chromatin accessibility, transcription factor binding, or splicing.

Comparing sequences before and after a variant

Input the original and variant sequences separately and compare their predictions. AlphaGenome is a representative model combining long sequences with multiple prediction tasks.

Connecting to cellular context

The same variant can have different effects in different tissues or cell states. Connecting sequence models with cell models requires aligning tissues, conditions, and what is measured.

  1. Sequence modelingDNA context

    Understanding coding and regulatory regions

  2. Multitask predictionAlphaGenome

    Predicting multiple molecular signals together

  3. Variant analysisControls and interventions

    Comparing predicted and experimental changes

Papers and learning resources

Noncoding region
A region of DNA that does not directly encode a protein; some of these regions help regulate genes.
Splicing
The RNA processing step that joins exons and removes introns. Different splicing patterns can produce different transcripts.

Try it yourself

Follow an official example to compare predicted tracks before and after a variant. Record the position, tissue, and output signal, explain which region changes, and look for corresponding experimental measurements.

Cells: Understand states and predict responses to interventions

Single-cell data record which genes are being expressed in each cell. Research moves from asking what a cell is and which state it is in to asking how it will respond to a gene change or a drug. Spatial omics also preserves the locations of cells within tissue.

Cell representations

Geneformer, scGPT, and UCE learn gene or cell representations from large amounts of single-cell data for tasks such as annotation and integration. Data processing needs to preserve biological differences while handling sequencing depth, noise, and batch effects.

Perturbations and virtual cells

Given a cellular context and a genetic or drug perturbation, predict changes in expression. Virtual cells aim to connect modalities, scales, and dynamic processes more broadly; perturbation response is one clearly defined research task within this goal.

Spatial relationships and biomanufacturing

FLAG predicts spatial gene expression from pathology images, focusing on relationships between genes and between tissue locations. Biomanufacturing also studies metabolic networks, expression burden, and resource allocation to help propose strain modifications.

  1. RepresentationGeneformer / scGPT / UCE

    Learning cell and gene representations

  2. RelationshipsscPRINT / FLAG

    Gene networks and spatial structure

  3. ResponseVirtual cells

    Predicting states after intervention

Read further

States, dynamics, and dataset construction

Many single-cell datasets are snapshots measured from different cells. Studying evolution over time requires sampling times, lineage information, or perturbation experiments. Data should record tissue, donor, batch, and measurement modality; perturbation data should also record the target, dose, time, and control.

Evaluating gene networks and spatial structure

Gene association networks can be compared with known regulation and intervention data. Spatial expression prediction should compare errors at individual locations, relationships between genes, and spatial distributions. Depending on the intended use, hold out new donors, new tissues, or unseen perturbations.

2026 · ICML

FLAG: Foundation model representation with Latent diffusion Alignment via Graph for spatial gene expression prediction

FLAG: Foundation model representation with Latent diffusion Alignment via Graph for spatial gene expression prediction · Figure 1
Figure excerpt from the paper · Figure 1 · Original paper

The problem

How can tissue images predict spatial gene expression while preserving structural relationships?

The key idea

Combine spatial graph encoding with gene foundation model representations, and use a diffusion process to model expression distributions.

Where to start

Start with input images, spatial adjacency, and expression outputs, then compare errors at individual locations with structural evaluations.

Papers and learning resources

Batch effect
Systematic variation in data caused by differences in experimental timing, equipment, or procedures.
Perturbation prediction
Predicting a cell response from interventions such as gene editing or drugs, together with the cellular context.

Try it yourself

Process a public single-cell dataset with Scanpy: first run quality control, PCA, and clustering, then compare pretrained representations. Color the results by cell type and by experimental batch to see what drives the groups.

EEG: From neural signals to state recognition and decoding

Electroencephalography (EEG) records how electrical signals at scalp electrodes change over time. AI can help recognize states, detect events, learn representations across participants, and connect signals with stimuli such as images that participants see.

Temporal and spatial representations

Electrode locations, sampling rate, filtering, and artifact processing together define the input. Models use frequency bands, changes over time, and relationships between channels.

Foundation models and transfer

After pretraining, the encoder can be frozen for linear probing or fine-tuned. Testing within the same participant and on new participants answers different questions. EEG-FM-Compass organizes these types of evaluation.

Visual decoding and generation

DreamDiffusion connects EEG representations to a pretrained image generation model and uses CLIP visual supervision. EEG-CLIP offers another alignment approach; compare how the two connect signals with images.

  1. Signal processingTime–frequency features

    Understanding data through channels and time

  2. 2023DreamDiffusion

    Signal alignment and conditional generation

  3. 2026EEG-FM-Compass

    Foundation models and transfer evaluation

Read further

Checking which information decoding uses

First fix the stimulus and participant splits, then compare correctly paired data, shuffled pairs, and results with EEG conditioning removed. This reveals how much the generated results depend on input signals, and what changes with new participants, sessions, or devices.

2023 · arXiv

DreamDiffusion: Generating High-Quality Images from Brain EEG Signals

DreamDiffusion: Generating High-Quality Images from Brain EEG Signals · Figure 1
Figure excerpt from the paper · Figure 1 · Original paper

The problem

How can EEG representations connect to an image generation model?

The key idea

Learn EEG representations through temporally masked signal modeling, combine them with CLIP visual supervision, and use the signals for conditional generation.

Where to start

Trace the role of each data source and supervision signal through the training stages, then examine participant splits and controls.

Papers and learning resources

Participant
A person who takes part in an experiment and provides neural signals.
Artifact
An unwanted signal introduced by sources such as eye movements, muscle activity, or equipment.

Try it yourself

Use MNE to inspect a public EEG recording. Plot the raw waveforms and power spectrum, then try event classification. Separate participants between training and testing, and compare traditional features with pretrained representations.

Medical imaging: Locate structures and measure motion and function

In CT, MRI, or pathology images, models can delineate organs and lesions, align different images, and calculate measurements such as volume from contours and motion. Imaging methods draw on computer vision, while their research questions come from specific examination and measurement needs.

Three-dimensional segmentation

3D U-Net and V-Net process volumetric data; UNet++ changes how features at multiple scales are fused; Swin UNETR introduces Transformers for context modeling. Compare how they preserve detail and use spatial relationships.

Registration and functional measurement

Registration estimates correspondences between images or different time frames. CMRINet jointly analyzes registration and segmentation in cardiac cine MRI, connecting structural identification to cardiac function quantification.

Longitudinal changes

Follow-up images help study disease progression and change. Record conditions such as sampling times, treatments, and missing follow-ups. Data from external centers and devices can test how broadly a method applies.

  1. 20163D U-Net / V-Net

    Volumetric segmentation

  2. 2018–2022UNet++ / Swin UNETR

    Feature fusion and global context

  3. 2025CMRINet

    Temporal structure and function analysis

2025 · arXiv

CMRINet: Joint Groupwise Registration and Segmentation for Cardiac Function Quantification from Cine-MRI

CMRINet: Joint Groupwise Registration and Segmentation for Cardiac Function Quantification from Cine-MRI · Figure 1
Figure excerpt from the paper · Figure 1 · Original paper

The problem

How can dynamic cardiac MRI provide both structure and motion information?

The key idea

Jointly handle groupwise registration and segmentation, connecting temporal image analysis to cardiac function quantification.

Where to start

Start with the temporal dimension of cine MRI, then examine how the two tasks work together and how functional measures are calculated.

Papers and learning resources

Voxel
A small volume element in a three-dimensional image, corresponding to a pixel in a two-dimensional image.
Dice / HD95
The former measures region overlap; the latter describes distances between contours. They capture different kinds of errors.

Try it yourself

Start with a segmentation example from MONAI and split the data by patient. Display the original image, annotation, and prediction side by side. Compare region overlap and boundary distance, and identify three typical error types.

Mathematics and algorithms: Search for proofs and discover better constructions

AI in mathematics can propose proof steps, add geometric constructions, or search for objects that satisfy constraints and for more efficient algorithms. A key advantage is that many candidates can be checked using formal proof systems, programs, or exact computation.

Proof search

AlphaGeometry combines auxiliary constructions proposed by a language model with symbolic reasoning. Formal proof approaches express propositions and steps in systems such as Lean, where a kernel checks the proof.

Mathematical constructions

FunSearch lets a language model propose programs, then runs an evaluation function to select useful candidates. The paper studies combinatorial problems including cap sets. The model searches for programs that produce constructions.

Algorithm discovery

AlphaEvolve combines code generation, evaluators, and evolutionary search to find better algorithms and mathematical constructions. Runtime, correctness, or objective values can provide feedback.

  1. 2023FunSearch

    Program generation + evaluation

  2. 2024AlphaGeometry / AlphaProof

    Construction search and proofs

  3. 2025AlphaEvolve

    Iterative algorithm optimization

2024 · Nature

Solving olympiad geometry without human demonstrations

Solving olympiad geometry without human demonstrations · Google DeepMind · AlphaGeometry 官方方法示意图
Figure excerpt from the paper · Google DeepMind · AlphaGeometry 官方方法示意图 · Original paper

The problem

How can auxiliary constructions and symbolic reasoning form a proof process?

The key idea

A symbolic engine first derives consequences from the given conditions. When needed, a language model proposes auxiliary constructions, and proof search continues.

Where to start

Trace the added constructions and the deductions they enable in a simple geometry problem, then examine how the training data are synthesized.

Papers and learning resources

Formal proof
A proof written with precise definitions and inference rules and checked by a proof assistant.
Evaluator
A program that executes a candidate and returns its correctness or objective score.

Try it yourself

Start by completing a few basic proofs in Lean. See which conditions you need to add between a step that looks correct and a proof the system accepts. Alternatively, write an evaluator for a bin-packing heuristic, let a model propose changes, and compare them on held-out instances.

Finance: From text and time-series data to predictions and decisions

Financial AI includes understanding financial reports and news, identifying risks, forecasting time series, constructing portfolios, and executing trades. Model inputs include structured data such as prices and trading volume, as well as text from announcements and reports.

Financial text

FinGPT offers an entry point to financial language models, data, and fine-tuning. Sentiment, event, and information extraction tasks help explain language modeling for this domain.

Quantitative prediction and backtesting

Qlib connects data processing, features, models, portfolios, and backtesting. A factor is a feature used to describe or predict market behavior; research needs to track when that information becomes available.

Decisions and automated research and development

FinRL studies sequential decision-making; RD-Agent-Quant lets agents iterate on factors and models. Evaluation considers returns, risk, turnover, transaction costs, and performance across different market periods.

  1. 2020Qlib

    A complete quantitative research workflow

  2. 2021–2023FinRL / FinGPT

    Decisions and financial language models

  3. 2025RD-Agent-Quant

    Joint factor and model optimization

2020 · arXiv

Qlib: An AI-oriented Quantitative Investment Platform

Qlib: An AI-oriented Quantitative Investment Platform · Figure 1
Figure excerpt from the paper · Figure 1 · Original paper

The problem

How do data, models, portfolios, and backtesting form a research workflow?

The key idea

Connect financial data processing and machine learning pipelines to strategy evaluation so that different models can be compared within a common workflow.

Where to start

Trace the data flow through the framework diagram, then identify the time split, labels, and backtesting configuration in an example.

Papers and learning resources

Backtesting
Moving through historical time to simulate the information available and the outcomes of decisions at each point.
Look-ahead bias
Using information in a prediction or decision that would not yet have been available at that time.

Try it yourself

Read a Qlib prediction and backtesting example, identifying the features, labels, training period, and test period. Compare a simple and a complex model using the same time split, and inspect results before and after transaction costs.

Graph learning and time series: Study relationships and change

Traffic, social interactions, transactions, and power grids all contain relationships and change over time. Graph learning asks who is connected to whom; time-series modeling asks how the past affects the future. The two can also be combined for spatiotemporal prediction.

Nodes, edges, and overall structure

Graph tasks can predict node classes, whether a relationship exists, or properties of an entire graph. Molecular graphs and social networks share methods such as message passing, but their labels and evaluation differ.

Trends, cycles, and forecasting

Time-series research distinguishes trends, cycles, and anomalies, and compares statistical methods with deep learning. Testing moves forward in time to match how forecasting is used in practice.

Spatiotemporal systems

Roads are connected in a road network, and traffic conditions change over time. Modeling needs to describe spatial dependencies while accounting for external events and time scales.

  1. RelationshipsGraphs and message passing

    Node, edge, and whole-graph tasks

  2. ChangeTrends and cycles

    Forecasting in temporal order

  3. CombinationSpatiotemporal graphs

    Dynamic systems such as traffic

Papers and learning resources

Message passing
Nodes aggregate information from neighbors and edges, updating their representations layer by layer.
Rolling forecasting
Repeatedly predicting a subsequent interval using only the data available as time moves forward.

Try it yourself

For a time series, compare the last observed value, a seasonal baseline, and a learned model using rolling forecasts. For a graph task, first sketch what nodes and edges represent, then decide how to split training and test data.

AutoML: Automatically select models and training configurations

Learning rate, tree depth, feature processing, and network architecture all affect results. AutoML studies how to automatically find better models and training plans for a given task and computation budget.

Hyperparameter optimization

Methods such as grid search, random search, and Bayesian optimization choose the next configuration. Optuna supports dynamic search spaces and pruning, stopping poorly performing trials early.

Model and pipeline selection

Automatically select data preprocessing, feature processing, models, and ensembles. For comparisons, keep data splits and budgets fixed and record the cost of the search itself.

Architecture search and meta-learning

Neural architecture search (NAS) looks for network structures; meta-learning uses experience from past tasks to help with new ones. Research can explore how to reduce search costs and transfer across tasks.

  1. ConfigurationHPO

    Hyperparameters such as learning rate

  2. PipelineModel and feature selection

    The complete learning pipeline

  3. Architecture and experienceNAS / Meta-learning

    Architecture search and knowledge across tasks

2019 · KDD

Optuna: A Next-generation Hyperparameter Optimization Framework

Optuna: A Next-generation Hyperparameter Optimization Framework · Akiba et al. · Figure 6
Figure excerpt from the paper · Akiba et al. · Figure 6 · Original paper

The problem

How can a limited training budget find a good configuration?

The key idea

Define the search space in code, then combine sampling and pruning to allocate the budget across trials.

Where to start

Start with the search space and trial loop, then examine how early stopping saves computation.

Papers and learning resources

Hyperparameter
A configuration set before training, such as learning rate, regularization strength, or number of layers.
Pruning
Using intermediate results to stop less promising training runs and save the budget for other trials.

Try it yourself

Give an existing classifier a budget of 20 training runs and compare random search with Optuna. Plot cumulative elapsed time against the best validation score so far. Finally, evaluate the selected configuration only on the held-out test set.

Automated research: Turn ideas into experiments, then use results to move forward

Research contains many executable steps: searching for information, proposing candidates, changing code, running experiments, analyzing results, and writing reports. Automated research systems try to connect these steps so that the last result informs the next experiment.

The experimental loop

autoresearch focuses on language model training experiments on a single GPU, letting an agent modify training code and compare results. Evaluation metrics determine whether changes are kept, and experiments have a fixed training budget.

From ideas to papers

AI Scientist brings experimental design, code execution, result analysis, and paper writing into one workflow. Version v2 further explores agentic tree search.

Research evidence and records

Each candidate is linked to its configuration, code version, run logs, and results. When reading about these systems, trace a conclusion back to the experiments that support it, and examine how failed candidates affect later choices.

  1. 2024AI Scientist

    Experiments, analysis, and writing

  2. 2025AI Scientist-v2

    Agentic tree search

  3. Training experimentsautoresearch

    Iteration within a fixed budget

2025 · arXiv

The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search · Yamada et al. · Figure 1
Figure excerpt from the paper · Yamada et al. · Figure 1 · Original paper

The problem

How can experimental results guide the next round of research?

The key idea

Use agentic tree search to organize experiments, connecting plans, execution, analysis, and paper writing.

Where to start

Start with one complete experimental branch and check whether the final claims correspond to actual results.

Papers and learning resources

Experiment budget
The permitted resources, such as training time, computing power, or number of calls.
Agentic tree search
Organizing experimental plans and subsequent changes as a tree, then using existing results to guide further exploration.

Try it yourself

Start with an existing small experiment. Fix the data, evaluation, and run budget, then let an agent propose and execute three modifications. Keep each change together with its logs and results, and explain which direction to choose next and why.

RSI: Can a system improve its own ability to improve?

Recursive Self-Improvement (RSI) asks a further question: after a system modifies itself, can it become better at making the next improvement? Research can target an agent's code, tools, or search strategy, and can also involve learning algorithms.

What is being modified?

AutoML usually searches configurations within a predefined space; automated research systems advance external research questions; self-improving systems also modify how they themselves work. When reading, first identify the code, strategies, or parameters that actually change.

Storing and selecting versions

Darwin Gödel Machine (DGM) lets a coding agent modify its own code, saves multiple versions, and uses coding tasks to evaluate them before exploring further. Its central loop creates, tests, retains, and modifies versions again.

Observing sustained improvement

Alongside scores on individual tasks, compare improvement curves over multiple rounds, total computation cost, and performance on new tasks. Design controls to distinguish a better next round caused by changes in system capability from one caused by spending more on search.

  1. Propose changesOwn code and tools

    Create a new version

  2. Run testsFixed tasks and budget

    Compare changes in capability

  3. Continue improvingVersion archive and selection

    Track effects over multiple rounds

2025 · arXiv

Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents

Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents · Figure 1
Figure excerpt from the paper · Figure 1 · Original paper

The problem

Can a coding agent improve its abilities by modifying its own code?

The key idea

Maintain multiple agent versions, generate self-modifications, and select candidates using coding tasks, creating an expanding version tree.

Where to start

Start with one specific code change, then examine version selection, evaluation tasks, and the computation budget.

Papers and learning resources

Self-modification
A system changes its own code, tools, or strategies. DGM mainly modifies agent code.
Recursive improvement
An improvement also changes the ability to improve further, forming subsequent iterations.

Try it yourself

Read the DGM version tree and choose a specific code change. Trace how it was proposed, how it was measured, and whether later versions inherited it. Then design a comparison between versions of a simple tool-using agent under a fixed budget.

Turn an interdisciplinary question into a small project

Once you have a question, connect the data, model and validation in a complete research workflow.

  1. Define the object of study. Is one sample a sequence, a patient, a cell or a market at a particular time? How is the target measured, and in what units?
  2. Preserve the structure of the domain. Molecules have symmetries, graphs have relationships, time series have an order, and medical images have voxel spacing. Which information do the representation, model and loss preserve?
  3. Decide what new conditions matter. New molecular scaffolds, protein families, donors, participants, patients and future periods call for different data splits.
  4. Establish a reference and analyse changes. Compare simple methods with the new model and locate its errors. Connect candidate designs to further calculations or experiments, recording how each result changes the next step.

For time-series foundations, Forecasting: Principles and Practice covers trends, forecasting and evaluation with examples in R. For Python projects, see tsai.

AI for Science resource index · AI4X resource directory · Research questions · Contact and collaboration · Experiments and records