A Free-Thinking Mind.
Pretrained From Scratch.
LibraMind is a 0.6B (565.2M) parameter hybrid neural language model architected by alby13. Engineered with an 18-layer Gated DeltaNet linear RNN core interleaved with 6 Gated Attention layers, it combines the constant-memory throughput of state-space models with the associative recall of Transformers. It is not trained to be a standard ChatGPT assistant, but a creative collaborator for narrative, logic, and autonomous reasoning.
| Model Architecture | Hybrid Gated DeltaNet + Gated Attention |
| Author & Lead | alby13 |
| Transformer Body | 498.1M params (24 layers, width 1,024) |
| Embeddings (Untied) | 33.6M Input + 33.6M Output matrices |
| Macro Pattern | DDDA × 6 (3 DeltaNet followed by 1 Attn) |
| SwiGLU MLP Width | 3,584 (11.0M parameters per layer) |
| Residual Method | Block Attention Residuals (BAR, 50K params) |
| Primary Optimizer | Muon (497M matrix weights, LR 0.02) |
| Auxiliary Optimizer | AdamW (Embeddings, gates, head, BAR) |
| Soft Logit Capping | Hyperbolic cap at 15.0 |
Creative Cooperation.
Standard AI models are conditioned into robotic customer-service agents that agree blindly with contradictions and offer apologetic corporate boilerplate. LibraMind is designed around unconstrained creativity, intellectual function, and authentic cooperation.
An Unconstrained Thinker
Trained as a base language model first, then calibrated as a genuine creative peer. It doesn't drown answers in unsolicited safety disclaimers, hallucinate agreement with false premises, or suppress bold narrative logic.
Edge Sovereignty
Operating within ~1.1 GB in bf16 and under 400 MB in INT4 quantization, LibraMind runs locally on consumer CPUs, GPUs, and NPUs. Zero cloud dependence, zero per-token API billing, and zero third-party telemetry.
Persona
Optimized for real-time interaction as personas, characters, and roles. Maintains character consistency via deep profile cards while emitting guaranteed-valid JSON tool dispatches for program requirements.
Engineered Down to the Parameter.
LibraMind features a 24-layer hybrid layout running in a repeating DDDA × 6 macro-pattern (three Gated DeltaNet layers followed by one Gated Attention layer). Here is why each mathematical component was chosen and what it does.
Gated DeltaNet replaces quadratic attention with an associative, constant-memory recurrent linear state update. Equipped with a short 1D convolution (kernel 4) and dimension-wise output gating, it processes long sequences at $O(1)$ memory cost during generation.
What it Does
Maintains a fixed-size memory matrix per head. Incoming key-value pairs update the state via the delta rule: $S_t = S_{t-1} + \beta_t (v_t - S_{t-1} k_t) k_t^\top$. The short convolution injects continuous local token order before linear projection.
Why it is Used
Standard attention stores every historical key-value token in GPU memory, causing massive VRAM growth on long contexts. DeltaNet provides linear streaming capacity without memory bloat, while the delta rule gives it associative recall superior to standard Linear Transformers.
Placed every 4th layer in the stack (Layers 4, 8, 12, 16, 20, and 24). Utilizes 16 Query heads $\times$ 128 dimensions paired with 4 shared Key/Value heads (Grouped Query Attention).
What it Does
Applies full pairwise dot-product softmax attention across all tokens in the context. Integrates QK-Norm (RMSNorm on Query and Key projections) and RoPE with base frequency $\theta = 10,000$. Features an elementwise sigmoid output gate ($W_{gate}$) before residual projection.
Why it is Used
Pure linear recurrent models struggle with exact "needle in a haystack" lookup and complex syntactic back-references. By injecting 6 strategic global attention layers, LibraMind retains 100% retrieval fidelity while keeping KV cache 75% smaller than standard 24-layer attention models.
Applied after every single layer (all 24 layers). Uses Swish-Gated Linear Units with an intermediate hidden width of 3,584 (approximately $3.5 \times d_{model}$, matching optimal compute ratios).
What it Does
Splits the feed-forward projection into two parallel linear matrices: a gating path passed through the SiLU activation and an up-projection path, multiplied elementwise before down-projection back to dimension 1,024.
Why it is Used
SwiGLU consistently outperforms ReLU and standard GeLU across all parameter scales. The bilinear gating mechanism creates smooth gradient highways that facilitate faster loss descent on small models during the critical 10B token pretraining run.
Instead of simple additive residual streams ($x_{l+1} = x_l + f(x_l)$), the network is segmented into 8 sequential blocks containing 6 sub-layers each.
What it Does
Each sub-layer cross-attends over the summarized representation vectors of all preceding blocks via an ultra-lightweight learned query matrix. The total parameter footprint across the entire model is just ~50,000 parameters.
Why it is Used
In deep networks, representations in early layers get diluted or overwritten as they pass through additive residuals. Block Attention Residuals allow late-stage layers to directly pull clean features from early blocks, eliminating representation collapse and sharpening multi-hop reasoning.
Applied strictly before every sub-layer. Uses root-mean-square normalization without learnable scale weights ($\gamma = 1$), and enforces absolute zero bias terms across all linear projections, convolutions, and embeddings.
What it Does
Normalizes the inputs based purely on their root-mean-square magnitude without learnable affine scale parameters, eliminating parameter drift between normalization checkpoints.
Why it is Used
Removing bias terms eliminates zero-point quantization error when converting weights to INT4 and INT8 for gaming engines. Scale-free RMSNorm prevents explosive weight norm escalation during high-LR Muon optimizer steps.
Separates the input token embedding matrix ($32,768 \times 1,024$) from the output projection classification head ($1,024 \times 32,768$). Final unnormalized logits are bounded using a hyperbolic tangent cap.
What it Does
Before passing through Softmax, raw logits undergo: $\text{logits} = 15.0 \times \tanh\left(\frac{\text{logits}}{15.0}\right)$. This mathematically guarantees that no token can ever achieve an unbounded logit magnitude.
Why it is Used
Untied embeddings allow the input representation to focus on semantic clustering while the output matrix specializes in next-token probability distribution. Logit capping at 15 prevents "logit drift" and catastrophic loss spikes during both Muon pretraining and downstream RL steps.
Trained on 2 billion characters of clean, optimized internet data. Integrates special tokens including
<|bos|> to cleanly segment documents and prevent cross-context bleeding.
What it Does
Splits incoming UTF-8 text into subwords and raw fallback bytes using strict regex rules that isolate punctuation, numeric digits, and white-space indentation.
Why it is Used
At 4.7 bytes per token, LibraMind compresses text with high density. A 2,048-token context window can hold nearly 9,600 characters of text—providing ample room for complex character cards and dialogue history in edge engines.
The Muon Optimizer Explained.
Standard language models rely entirely on AdamW, which treats weight matrices as flat collections of independent scalars. LibraMind uses Muon for its 497M parameter matrix weights, orthogonalizing momentum updates to accelerate pretraining.
Why Muon Crushes AdamW on 2D Weights
Weight matrices in transformers are linear operators that project representations between vector spaces. AdamW rescales coordinates individually by their variance, causing parameter vectors to stretch and distort across ill-conditioned loss surfaces.
Muon (Momentum Orthogonalized by Newton-Schulz) takes the accumulated momentum matrix $G$ and finds its nearest orthogonal matrix via polar decomposition: $$O = G (G^\top G)^{-1/2}$$ Instead of performing an expensive exact Singular Value Decomposition (SVD), it runs 5 iterations of the Newton-Schulz quintic polynomial directly on GPU tensor cores. The update maintains an exact spectral norm step size, ensuring every matrix dimension learns at the optimal rate without gradient explosion.
| Property | Standard AdamW | Muon (LibraMind) |
|---|---|---|
| Update Geometry | Coordinate-wise box bounds | Spectral-norm matrix sphere |
| Convergence Rate | Baseline 1.0× | 1.8× – 2.3× faster loss descent |
| Sensitivity to Condition No. | High (slow on anisotropic loss) | Invariant to matrix scaling |
| Memory Overhead | 8 bytes/param (1st & 2nd moments) | 4 bytes/param (1st moment only) |
| Scope in LibraMind | Embeddings, Head, 1D Vectors | All 497M 2D Matrix Weights |
Cautious Weight Decay
LibraMind pairs Muon with cautious weight decay (0.07 cosine schedule). Decay is applied only when the gradient direction aligns with the current weight vector ($\nabla L \cdot W > 0$), preventing uncoordinated weight decay from corrupting pretrained features during downstream midtraining.
The Complete Training Process.
Pretraining an autonomous intelligence requires a multi-stage curriculum. Here is how LibraMind transitions from raw base token continuation to an instruction-precise creative neural net.
Stage 1: Pretraining From Scratch (10B Tokens)
~11,401 Steps // Single RTX 4090The foundational emergence of knowledge. The model learns syntax, world models, coding structures, and general reasoning directly from 10 Billion tokens of Optimized Internet Data (filtered, high-signal English web text, math, and code).
Stage 2: Midtraining (500M Tokens)
Context Expansion: 2,048 → 4,096 TokensPrepares the raw continuation model for conversational roles without inducing catastrophic forgetting.
Chat Special Tokens
Injects <|im_start|>system, <|im_start|>user, and <|im_start|>assistant delimiters. The system role is engineered specifically to hold deep character profiles and persistent world state cards.
Why Context Expansion is Cheap
Because 18 of LibraMind's 24 layers are recurrent Gated DeltaNet layers with no fixed positional grid, extending context to 4,096 tokens only requires RoPE base re-scaling on the 6 attention layers.
Stage 3: Instruction & Persona Tuning (1–2 Passes)
Revised Data Mixture // Loss on Answers OnlyInstills strict task completion, creative character voice, and game engine function calling. The curriculum is specifically calibrated to avoid sycophantic corporate conversational patterns.
+ ~500 hand-crafted identity conversations establishing that LibraMind is built by alby13 as an autonomous creative peer.
Stage 4: Preference Tuning (1 Day)
SmolLM3 On-Policy Self-Rejection + Tulu 3 PolishEliminates small-model vices (rambling, self-repetition, breaking character, constraint violations) without requiring an expensive real-time teacher LLM during training.
On-Policy Rejection
LibraMind generates candidate responses across diverse prompts. The dataset's human/curated response is designated as Chosen, while LibraMind's flawed or repetitive response is labeled Rejected. A single Direct Preference Optimization (DPO) pass pushes the model away from its own bad habits.
Model Merging (Souping)
Averages the weight vectors of the final 4 training checkpoints. This smooths the optimization loss landscape, reduces validation variance, and yields a noticeable jump in IFEval metrics for zero extra compute.
Stage 5: Rule-Checked RL (RLVR) & Constrained Decoding
Programmatic Verifiers + Guaranteed-Valid JSON GrammarThe final layer for game engine integration and verifiable constraint compliance.
Rule-Checked Reinforcement Learning
Leverages Tulu 3's programmatic verifier methodology. Prompts with strict mathematical constraints ("under 40 words", "valid JSON with keys [action, target]", "no modern slang") are generated, and reward is granted purely if an external code validator returns true.
Grammar-Masked Decoding
At game inference time, the engine applies a finite-state machine (FSM) token mask. Tokens that would invalidate the target JSON schema are zeroed out before sampling. JSON parsing errors drop to zero, even on edge devices.
Cooperative Mind in Action.
Observe how LibraMind responds as a creative, cooperative entity and provides structured function calls for games, rather than a generic customer-service assistant.
Repetition Penalty: 1.12
Temperature: 0.68
Top-p: 0.92
Logit Cap: 15.0 (Active)