How TRACED learns, chunk by chunk

The real spdal implementation on small two-dimensional streams

TRACED learns from a stream that arrives in chunks. Each chunk is absorbed into hyperellipsoidal neurons (a centre, principal axes, and a width along each axis) and then discarded. During learning a neuron can be

Each neuron also keeps a smoothed displacement (how its centre has been moving) and expansion (how its widths have been changing). When a test point lies outside every neuron, the trend rule shifts and stretches the two nearest neurons by these trends before deciding, so a class that has been moving towards the point can claim it.

Everything here comes from spdal 0.4.1 (TRACED.partial_fit and TRACED.predict), unmodified. Learning events are counted by wrapping TRACED’s own methods.

Setup. Gaussian classes, 20 chunks of 50 points, processed test-then-train, with r = 1, δ = 4, α = 0.5 (the default) and β = 0.5. In our runs, changing α altered the length of the displacement arrows but not the accuracy, so it is fixed here. Five streams:

  • Stationary: two classes that stay put.
  • B moves towards A / B moves past A: both classes start from the same distribution for four chunks. Class B then shifts gradually towards a target point (slow start, faster in the middle, slowing down on arrival by chunk 16) and stays there, either close beside A or on A’s far side.
  • A new class appears: a third class starts arriving at chunk 8.
  • With outliers: 5% of the points are replaced by uniform noise with random labels.

In two dimensions the coincident-region rule (D4) would project onto a single axis, so it is switched off (reduce_dims=0); see the HDFR-B demo for D4 in 20 dimensions.

How the moving streams were chosen. We tried five target points, two arrival times, two motion profiles and two chunk sizes, each with 24 hyperparameter settings and 3 random seeds (2,880 runs). “Past A” is where the trend rule helped most; averaged over all of these streams and settings it helped by about +0.3 points.

Run TRACED on the five streams
import itertools, warnings
import numpy as np
from spdal import TRACED

warnings.filterwarnings("ignore")

class CountingTRACED(TRACED):
    """TRACED with counters around its own learning steps. Behaviour is unchanged."""
    def reset_events(self):
        self.events = {"created": 0, "updated": 0, "merged": 0}
    def create_new_neuron(self, X, y):
        self.events["created"] += 1
        return super().create_new_neuron(X, y)
    def update_parameter(self, neuron, alpha, X, Y, Y_index):
        self.events["updated"] += 1
        return super().update_parameter(neuron, alpha, X, Y, Y_index)
    def merge_neuron(self, alpha, y):
        merged, alpha = super().merge_neuron(alpha, y)
        self.events["merged"] += int(merged)
        return merged, alpha

N_CHUNKS, CHUNK, START, ARRIVE, APPEAR, OUTLIERS = 20, 50, 4, 16, 8, 0.05
SD = np.array([0.5, 0.35])
CENTRES = np.array([[0.0, 0.0], [3.0, 0.0], [1.5, 2.2]])            # class A, class B, class C (new)
TARGETS = {"towards": np.array([1.0, 0.0]), "past": np.array([-1.5, 1.2])}
LO, HI, NX, NY = np.array([-3.5, -2.0]), np.array([4.5, 3.5]), 56, 39
SCENARIOS = ["stationary", "towards", "past", "new", "outliers"]

def centre_b(k, scenario):
    """Class B stays at its first distribution, then moves gradually to its target (ease in and out)."""
    if scenario not in TARGETS:
        return CENTRES[1]
    t = np.clip((k - START) / (ARRIVE - START), 0, 1)
    t = t * t * (3 - 2 * t)
    return CENTRES[1] + t * (TARGETS[scenario] - CENTRES[1])

def make_stream(scenario, seed):
    rng = np.random.default_rng(seed)
    chunks = []
    for k in range(N_CHUNKS):
        n_classes = 3 if scenario == "new" and k >= APPEAR else 2
        y = rng.integers(0, n_classes, CHUNK)
        centres = CENTRES.copy()
        centres[1] = centre_b(k, scenario)
        X = centres[y] + rng.normal(0, 1, (CHUNK, 2)) * SD
        if scenario == "outliers":
            noisy = rng.random(CHUNK) < OUTLIERS
            X[noisy] = rng.uniform(LO, HI, (noisy.sum(), 2))
        chunks.append((X, y))
    return chunks

gx, gy = np.linspace(LO[0], HI[0], NX), np.linspace(LO[1], HI[1], NY)
GRID = np.array([[a, b] for b in gy for a in gx])

def run(scenario, seed, record=False, alpha=0.5):
    chunks = make_stream(scenario, seed)
    classes = [0, 1, 2] if scenario == "new" else [0, 1]
    model = CountingTRACED(r=1, delta=4, alpha=alpha, beta=0.5, reduce_dims=0, method="none")
    steps, acc = [], {"none": [], "outside": []}
    for k in range(N_CHUNKS - 1):
        model.reset_events()
        model.partial_fit(*chunks[k], classes=classes)
        X_next, y_next = chunks[k + 1]
        pred = {}
        for method in acc:
            model.method = method
            pred[method] = model.predict(X_next).astype(int)
            acc[method].append(float(np.mean(pred[method] == y_next)))
        if not record:
            continue
        maps = {}
        for method in acc:
            model.method = method
            maps[method] = model.predict(GRID).astype(int)
        status = np.where((pred["none"] != y_next) & (pred["outside"] == y_next), 1,          # fixed by the rule
                 np.where((pred["none"] == y_next) & (pred["outside"] != y_next), -1, 0))     # broken by the rule
        steps.append({
            "neurons": [[*np.round(n["center"], 3), *np.round(n["eig_component"].ravel(), 3),
                         *np.round(n["width"], 3), int(n["y"]), int(n["n"]), *np.round(n["displacement"], 3)]
                        for n in model.neuron_list],
            "decision": "".join(map(str, maps["none"])),
            "changed": np.flatnonzero(maps["none"] != maps["outside"]).tolist(),
            "test": [[round(float(a), 3), round(float(b), 3), int(c), int(s)]
                     for (a, b), c, s in zip(X_next, y_next, status)],
            "learned": [[round(float(a), 3), round(float(b), 3), int(c)] for (a, b), c in zip(*chunks[k])],
            "target": centre_b(k + 1, scenario).round(3).tolist() if scenario in TARGETS else None,
            "events": dict(model.events),
            "acc": {m: acc[m][-1] for m in acc},
        })
    return steps, {m: float(np.mean(v)) for m, v in acc.items()}, acc

runs, summary = {}, {}
for scenario in SCENARIOS:
    key = scenario
    runs[key] = run(scenario, seed=0, record=True)[0]
    per_seed = [run(scenario, seed=s) for s in range(5)]
    gains = [p[1]["outside"] - p[1]["none"] for p in per_seed]
    summary[key] = {"none": float(np.mean([p[1]["none"] for p in per_seed])),
                    "outside": float(np.mean([p[1]["outside"] for p in per_seed])),
                    "gain_min": float(min(gains)), "gain_max": float(max(gains)),
                    # accuracy on the first chunk that contains the new class, and on the one after it
                    "appear": [float(np.mean([p[2]["none"][APPEAR - 1] for p in per_seed])),
                               float(np.mean([p[2]["none"][APPEAR] for p in per_seed]))]}

ojs_define(tr_runs=runs, tr_summary=summary,
           tr_meta={"nx": NX, "ny": NY, "lo": LO.tolist(), "hi": HI.tolist(), "n_steps": N_CHUNKS - 1,
                    "N0": TRACED().N0, "appear": APPEAR})

Explore

Top: the background is TRACED’s decision without the trend rule; small squares mark places where the trend rule gives a different class. Faint dots are the chunk just learned; solid dots are the next chunk, which the model is tested on. A green ring marks a test point the trend rule gets right that the plain rule got wrong; a red cross marks the opposite. Arrows show each neuron’s displacement, drawn 4× longer so they are visible. Only neurons with at least N₀ = 3 samples are drawn; smaller ones are ignored by predict and hidden here. In the moving streams the + marks where the centre of class B’s distribution actually is. Bottom: accuracy on each next chunk.

What to look for

  • Learning is local. Press Play in the stationary stream. Most samples update the neurons that already cover them; samples that no neuron can capture start new neurons, and intersecting neurons of the same class merge.
  • A new class is learned from a single chunk. When the third class first shows up, the model has never seen it, so accuracy on that chunk drops (about 68% on average over five seeds). After training on that one chunk, accuracy is back to 100% on the next chunk in every seed. There is no retraining from scratch: the new class simply gets neurons of its own, next to the existing ones.
  • Outliers start neurons of their own, and some of them stick. A noise point that no neuron can capture starts a new neuron. About half of these never reach N₀ samples, so predict ignores them (and they are not drawn); the rest gather enough outliers of the same label to pass N₀ (over five random streams, 12 such neurons stayed below N₀ and 11 passed it). Play the stream to the end: a large orange neuron built from outliers sits in the empty upper left and claims that space for class B. Accuracy still averages about 96%, most likely because few real points fall there. The trend rule does slightly worse than the plain rule here (about −0.6 points on average): outliers often lie outside every neuron, and shifting the neurons by their trend sometimes hands them to the wrong class.
  • Class B’s neurons follow it, but slowly. In the moving streams, B’s neurons absorb the new samples and their centres drift after it (the + marks where B actually is). A neuron’s centre is the average of everything it has absorbed, so its displacement per update is small compared with how far B moves (up to 0.25 per chunk towards A and 0.58 per chunk past A). Merging two neurons also resets their trend (Algorithm 2 of the paper).
  • Moving towards A: the trend rule changes almost nothing (+0.1 points at most).
  • Moving past A: the rule helps a little, and reliably. When B crosses over A, accuracy without the rule drops to about 62% in the worst chunk. The rule rescues some of those points (green rings): about +0.8 points on average, and positive in every seed.
  • Old neurons stay behind. TRACED does not forget, so neurons learned where B used to be remain there after it has moved on.