Abstract
The full technical report on a latent-workspace language model: a frozen Qwen3-0.6B decoder wrapped in 73M–132M trainable sidecars implementing a root-preserving expert graph, isolated parallel reasoning lanes, and a private diffusion canvas behind a calibrated publication rule. Ten preregistered prototype experiments falsified the architecture at this scale.
Routing passed its load gates on one seed of two but failed causal liveness on 10 of 10 seed-runs. Verified fan-in lost to self-consistency by roughly 30 points, traced to a beginning-of-sequence id that the pinned tokenizer aliases to end-of-text.
A dedicated repair experiment fixed that channel — zero turn-scaffold emissions across 640 audited rows — and then measured the workspace directly. Ablating the read-out's 34 memory positions changed graded accuracy by 0.000 on one seed and improved it by 0.022 on the other. Parity was reached by irrelevance.
Key findings
- Routing failed causal liveness on 10 of 10 seed-runs.
- Verified fan-in lost to self-consistency by roughly 30 points.
- Ablating all 34 workspace memory positions changed accuracy by 0.000 and 0.022.
- The matched plain-LoRA control never ran, so no capability claim is made.
Read it here
Cite this paper
The DOI below is the concept DOI and always resolves to the latest version. To pin this exact revision, cite 10.5281/zenodo.22538272 instead.
@misc{yaddanapudi2026hierarchical,
author = {Karthik Yaddanapudi},
title = {The Hierarchical Latent Workspace Model: Root-Preserving Variable-Depth Expertise for Parallel Diffusion Reasoning},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.22538271},
url = {https://doi.org/10.5281/zenodo.22538271}
}
All papers ↗