OpenArchEvo

arXiv 2609.40258LLM-guided evolution · native spiking networks

Architecture search beyond the search space.

OpenArchEvo makes an LLM the mutation operator of a population-based evolutionary search: the model writes complete network programs; a surrogate and a three-view novelty measure decide which survive. Its spiking discovery NeuroGate reaches 26.4 WikiText-103 PPL, ahead of the ANN DeltaNet (27.5), at 31.7× lower estimated arithmetic energy than a dense Transformer.

Ruoyu Zhao*, Jiaqi Wu*, Chenyu Zhu, Zhichao Lu *equal contribution

Large Language Model-Guided Evolutionary Discovery of Native Neural Architectures for Spiking Sequence Modeling

Search trajectory · real data

Ten iterations, 171 trained architectures

Each dot is an architecture trained on both proxy tasks: 20 initial experts and 151 found by the search. Press play to watch the cumulative Pareto front advance toward lower perplexity and higher accuracy. Click a dot to inspect it; double-click to open it.

About the dataSearch-proxy results of all 171 trained architectures
  • WT2 PPL: WikiText-2 perplexity after 30 training epochs. ListOps acc.: accuracy after 25 epochs on ⅓ of the ListOps training set.
  • Fitness combines the two with the paper formula (App. D.2); the SpikingDeltaNet expert (58.28 PPL, 53.0%) scores 0.500.
  • Runs above 200 PPL (diverged; 9 in total) are kept and drawn in the band on the right of the full-range view.

Method

Two loops: explore cheaply, train selectively

An inner loop generates thousands of feasible programs per iteration and ranks them with a surrogate. An outer loop spends real GPU time on at most 16 of them, chosen for predicted performance and novelty, and feeds the measured results back.

OpenArchEvo pipeline: inner loop (initialization and variation, three-view representation, screening and prediction, multi-island evolution) and outer loop (shortlisting, NSGA-II selection, real training, archive update)
Overview (paper Fig. 2). Click to enlarge.

INNER LLM-driven program evolution

  1. Seed 10 islands with strong and diverse programs from the trained archive.
  2. Mutate: the LLM rewrites or recombines parent programs sampled within or across islands.
  3. Screen out broken, non-causal and near-duplicate programs; the surrogate predicts the rest.
  4. Survive: islands keep their best-predicted programs and weak islands are reseeded, until ~1,080 are accepted.

OUTER Selective real training

  1. Select at most 16 candidates by NSGA-II on predicted fitness and novelty.
  2. Train each on the WikiText-2 and ListOps proxies.
  3. Archive the measured results and refresh the surrogate.
  4. Repeat for 10 iterations: 151 new architectures.

Representation · core contribution

How do you compare architectures nobody has trained yet?

When candidates are long architecture programs, the hard question is not how you search (evolution or a multi-agent system) but how you describe a network so that you can compare candidates and predict their quality. OpenArchEvo describes every program at three levels.

Three-view SNN architecture representation: (a) code, compared by n-grams, syntax and dataflow; (b) design rationale, embedded by a text model; (c) behavioural fingerprint, a 21-D vector of module statistics, structural statistics and initialization-time probes such as SWSP and FireRate
Paper Fig. 3. Click to enlarge.
  1. 01 · Code

    What it implements

    Operations and dataflow, compared with CodeBLEU.

  2. 02 · Design rationale

    What it intends

    The LLM’s stated idea, compared by text embedding.

  3. 03 · Behavioural fingerprint

    How it behaves

    21 numbers measured before training, compared by cosine.

Novelty is the mean three-view distance to a reference set. The fingerprint alone also feeds the surrogate and the near-duplicate filter.

d = 1 − ⅓(scode + srat + sfp)
𝒩(a) = meana′∈ℛ d(a, a′)

Inside the fingerprint

Seven probes on the untrained network, plus 14 statistics read from its modules and execution trace.

SWSP and firing rate, read from spikes. Four probe inputs drive the same six neurons. Firing rate counts how often they spike. SWSP counts how many distinct patterns a neuron–time cell shows across the inputs: high when responses depend on the input. Illustrative, not real data.

Two architectures, two fingerprints

Logged fingerprints of trained architectures. Each bar is scaled with the fixed bounds used for similarity: the dashed ticks are the 5th and 95th percentiles of the calibration set.

From fingerprint to predicted quality

Why TabPFN-2.5: the archive holds only tens to hundreds of trained architectures, and a pretrained tabular model predicts from them in context, with no surrogate training in each iteration.

Inside one iteration · real data

What the inner loop explored before anything was trained

Each grey dot is a program the LLM wrote and the surrogate scored; the coloured dots are the few that were then trained (hover to compare prediction with reality).

Results

Final Pareto front

After iteration 10, no other trained architecture beats these on both search proxies: lower WT2 perplexity and higher ListOps accuracy.

Paper resultsFull-scale retraining on WikiText-103 and Long Range Arena (Table 1) and spiking ablations (Table 2)

Are the evolved elements really spike-native?

Ablations on WT2 under the fixed search protocol (paper Table 2). Removing HomeoResSSM’s spike-driven control costs 6.9 PPL; restoring DeltaNet’s SiLU hurts NeuroGate; LoopMem’s state-norm feedback helps the SNN but slightly hurts a non-spiking version.

Architecture explorer

Every trained architecture, with code and diagram

Open any architecture for its diagram, code, rationale and metrics. Browse all diagrams →

Cite

BibTeX