ALER: Adaptive Learnable Experience Rewriting
for Reinforcement Learning

Oleg Shchendrigin, Egor Cherepanov, Alexey K. Kovalev, Aleksandr I. Panov

Arxiv preprint

TL;DR. In partially observable RL, a later observation can make stored information obsolete or change what it implies for the next decision. We formalize rewriting and experience fusion next to retention, introduce the Rune Mazes environments that test them, and propose ALER, an agent that pairs an LSTM with an independently addressed slot memory and a learned fusion gate.

Three memory updates

Consider a T-Maze in which the agent sees a cue at the start of a corridor and must turn toward the side it indicates at the junction. If a later corridor shows a new cue, the old cue becomes obsolete and the agent has to rewrite it. If the corridor contains a rune that inverts the cue, the agent has to turn the other way, which it can do only if it still holds the cue and applies the inversion. Each observation therefore maps the decision-relevant content to itself, to a new value, or to a different function of it.

UpdateWhat the observation does to the targetExample
Retentionkeeps it unchangedcorridor step, NoOp rune
Rewritingsets it to a value independent of the old onenew cue in Endless T-Maze
Experience fusiontransforms it by a rule that the observation specifiesInvert rune

For tasks composed of such updates, we count the memory states that a solution needs. The count separates an online solution, which folds every rune into the stored cue, from a deferred solution, which keeps the cue unchanged and tracks the runes apart from it. Several baselines reach their lowest success rates on the compositions that need more states.

Method

ALER architecture
ALER architecture. The shared LSTM state ht queries the episodic memory (Read), a learned gate combines the retrieved vector rt with ht into zt for the actor and the critic (Fusion), and an independently addressed Gumbel-Softmax write stores the write candidate xt (Write). PPO trains all parts jointly.

Environments

Rune Mazes are three environments in which rune observations invert, cancel, reset, or repeat updates of a hidden cue. Together with Endless T-Maze, in which each of n consecutive corridors shows a new cue, they cover all three update classes under vector and pixel observations.

Rune Multi-Corridor and Rune T-Maze
Vector Rune Mazes. Left: Rune Multi-Corridor, in which a multi-bit cue indicates one of K ∈ {4, 8, 16} branches and an Invert rune changes it. Right: Rune T-Maze, in which the agent must invert the initial cue after an Invert rune and ignore a NoOp rune.
Rune MiniGrid Memory rollout
Rune MiniGrid Memory. One pixel-based episode with RepeatPrev + SkipNext + Invert, from the cue room (left) to the junction (right). Each rune fires once on contact.
Endless T-Maze
Endless T-Maze. Every further corridor shows a new cue and requires one rewrite.
Invert swaps the success and failure targets.
SkipNext cancels the next rune.
ResetRules restores the original targets.
RepeatPrev re-applies the last rune that took effect.
NoOp leaves the rules unchanged.

Results

We compare ALER with seven baselines: PPO-MLP, PPO-LSTM, GTrXL, S5, FFM, SHM, and DNC. ALER reaches a success rate of at least 0.82 in all sixteen Endless T-Maze configurations and at least 0.99 on all five Rune T-Maze compositions, and it leads PPO-LSTM in eight of ten Rune MiniGrid Memory configurations.

Rune T-Maze

CompositionQ*PPO-MLPPPO-LSTMGTrXLS5FFMSHMDNCALER
Invert20.42±0.110.84±0.041.00±0.001.00±0.001.00±0.000.72±0.091.00±0.001.00±0.00
NoOp + Invert20.51±0.000.78±0.001.00±0.001.00±0.001.00±0.000.74±0.060.48±0.071.00±0.00
SkipNext + Invert40.51±0.011.00±0.001.00±0.001.00±0.000.65±0.080.53±0.000.56±0.021.00±0.00
3 Invert + ResetRules40.42±0.110.81±0.050.99±0.000.99±0.000.59±0.030.51±0.020.48±0.110.99±0.00
RepeatPrev + SkipNext + Invert100.41±0.100.83±0.020.92±0.080.99±0.010.62±0.010.55±0.010.50±0.000.99±0.01

Success rate on 100 validation episodes, mean ± SEM over five training seeds. Q* is the number of target states that the composition requires. Bold marks the best mean in each row.

Horizon extrapolation
Endless T-Maze horizon extrapolation. Agents trained on 10 corridors of length 10 and evaluated on 10 corridors of lengths 10 to 100. ALER stays at 1.00 up to length 30, and PPO-LSTM falls below 0.5 at length 20.

Rune MiniGrid Memory (pixels)

Fixed lengthRandom length
RunesPPO-LSTMALERPPO-LSTMALER
I0.80±0.030.69±0.090.93±0.070.94±0.06
N+I0.94±0.060.77±0.030.67±0.070.92±0.08
S+I0.51±0.030.59±0.010.62±0.070.63±0.05
3I+R0.53±0.040.57±0.010.53±0.010.57±0.01
P+S+I0.49±0.030.60±0.000.56±0.000.69±0.06

Mean ± SEM over three training seeds, 30M steps. I = Invert, N = NoOp, S = SkipNext, R = ResetRules, P = RepeatPrev. Both agents share the CNN encoder, LSTM size, and PPO settings.

What ALER stores

Linear probes show that the learned memory follows the deferred solution. On Rune T-Maze with RepeatPrev + SkipNext + Invert, the slot memory decodes the initial cue and the recurrent state decodes the rune flags. In Endless T-Maze, the first observation of a new cue raises its decoding in memory from 0.52 before the write to 0.99 after it, so one write stores the new cue. Against constant gates with frozen weights, the learned gate has the highest mean success rate on all three probed tasks.

Probe balanced accuracy
Linear probes on Rune T-Maze. Balanced accuracy for the initial cue, the rune flags, and the required turn, decoded from each internal representation of ALER. Chance is 0.5.