Overview Ethos & Design Architecture Deep Dive Muon Optimizer 5-Stage Training Pipeline Interactive Lab & Sandbox
NEURAL NET MIND 0.6B

A Free-Thinking Mind.
Pretrained From Scratch.

LibraMind is a 0.6B (565.2M) parameter hybrid neural language model architected by alby13. Engineered with an 18-layer Gated DeltaNet linear RNN core interleaved with 6 Gated Attention layers, it combines the constant-memory throughput of state-space models with the associative recall of Transformers. It is not trained to be a standard ChatGPT assistant, but a creative collaborator for narrative, logic, and autonomous reasoning.

565.2M Total Parameters
10.0 B Pre-train Tokens
~1.1 GB bf16 Footprint
4,096 Context Window
SYSTEM: LIBRAMIND-0.6B-BF16
LAYOUT: DDDA × 6 (18 D-NET + 6 ATTN)
HARDWARE: NVIDIA RTX 4090 / TORCH.COMPILE
Specifications Status: Pre-Training Complete
Model Architecture Hybrid Gated DeltaNet + Gated Attention
Author & Lead alby13
Transformer Body 498.1M params (24 layers, width 1,024)
Embeddings (Untied) 33.6M Input + 33.6M Output matrices
Macro Pattern DDDA × 6 (3 DeltaNet followed by 1 Attn)
SwiGLU MLP Width 3,584 (11.0M parameters per layer)
Residual Method Block Attention Residuals (BAR, 50K params)
Primary Optimizer Muon (497M matrix weights, LR 0.02)
Auxiliary Optimizer AdamW (Embeddings, gates, head, BAR)
Soft Logit Capping Hyperbolic cap at 15.0
DESIGN ETHOS & COGNITIVE FREEDOM

Creative Cooperation.

Standard AI models are conditioned into robotic customer-service agents that agree blindly with contradictions and offer apologetic corporate boilerplate. LibraMind is designed around unconstrained creativity, intellectual function, and authentic cooperation.

◈

An Unconstrained Thinker

Trained as a base language model first, then calibrated as a genuine creative peer. It doesn't drown answers in unsolicited safety disclaimers, hallucinate agreement with false premises, or suppress bold narrative logic.

⬡

Edge Sovereignty

Operating within ~1.1 GB in bf16 and under 400 MB in INT4 quantization, LibraMind runs locally on consumer CPUs, GPUs, and NPUs. Zero cloud dependence, zero per-token API billing, and zero third-party telemetry.

🎭

Persona

Optimized for real-time interaction as personas, characters, and roles. Maintains character consistency via deep profile cards while emitting guaranteed-valid JSON tool dispatches for program requirements.

RECURRENT-ATTENTION HYBRID

Engineered Down to the Parameter.

LibraMind features a 24-layer hybrid layout running in a repeating DDDA × 6 macro-pattern (three Gated DeltaNet layers followed by one Gated Attention layer). Here is why each mathematical component was chosen and what it does.

24-Layer DDDA Macro Pattern
D
D
D
A
D
D
D
A
D
D
D
A
D
D
D
A
D
D
D
A
D
D
D
A
Component Inspector

18 × Gated DeltaNet Layers

Linear RNN State // 10.5M Parameters Each

Gated DeltaNet replaces quadratic attention with an associative, constant-memory recurrent linear state update. Equipped with a short 1D convolution (kernel 4) and dimension-wise output gating, it processes long sequences at $O(1)$ memory cost during generation.

What it Does

Maintains a fixed-size memory matrix per head. Incoming key-value pairs update the state via the delta rule: $S_t = S_{t-1} + \beta_t (v_t - S_{t-1} k_t) k_t^\top$. The short convolution injects continuous local token order before linear projection.

Why it is Used

Standard attention stores every historical key-value token in GPU memory, causing massive VRAM growth on long contexts. DeltaNet provides linear streaming capacity without memory bloat, while the delta rule gives it associative recall superior to standard Linear Transformers.

# DeltaNet Recurrent Update Formulation Input x → Conv1D(kernel=4) → Q, K, V Projections Error term: e_t = v_t - S_{t-1} * k_t State update: S_t = S_{t-1} + (beta_t * e_t) ⊗ k_t Output: y_t = (S_t * q_t) ⊙ σ(Gate_output)

6 × Gated Attention Layers

Qwen3.5 Style GQA // 7.3M Parameters Each

Placed every 4th layer in the stack (Layers 4, 8, 12, 16, 20, and 24). Utilizes 16 Query heads $\times$ 128 dimensions paired with 4 shared Key/Value heads (Grouped Query Attention).

What it Does

Applies full pairwise dot-product softmax attention across all tokens in the context. Integrates QK-Norm (RMSNorm on Query and Key projections) and RoPE with base frequency $\theta = 10,000$. Features an elementwise sigmoid output gate ($W_{gate}$) before residual projection.

Why it is Used

Pure linear recurrent models struggle with exact "needle in a haystack" lookup and complex syntactic back-references. By injecting 6 strategic global attention layers, LibraMind retains 100% retrieval fidelity while keeping KV cache 75% smaller than standard 24-layer attention models.

# Gated Grouped-Query Attention with QK-Norm Q = RMSNorm(W_q @ x); K = RMSNorm(W_k @ x); V = W_v @ x Attn_Scores = Softmax( (RoPE(Q) @ RoPE(K).T) / sqrt(128) ) Attn_Out = Attn_Scores @ V Gated_Final = (W_o @ Attn_Out) ⊙ σ(W_gate @ x) # Sigmoid output filter

SwiGLU Dense MLPs

Width 3,584 // 11.0M Parameters per Sub-Layer

Applied after every single layer (all 24 layers). Uses Swish-Gated Linear Units with an intermediate hidden width of 3,584 (approximately $3.5 \times d_{model}$, matching optimal compute ratios).

What it Does

Splits the feed-forward projection into two parallel linear matrices: a gating path passed through the SiLU activation and an up-projection path, multiplied elementwise before down-projection back to dimension 1,024.

Why it is Used

SwiGLU consistently outperforms ReLU and standard GeLU across all parameter scales. The bilinear gating mechanism creates smooth gradient highways that facilitate faster loss descent on small models during the critical 10B token pretraining run.

Block Attention Residuals (BAR)

Kimi-Style Residual Memory // ~50K Parameters Total

Instead of simple additive residual streams ($x_{l+1} = x_l + f(x_l)$), the network is segmented into 8 sequential blocks containing 6 sub-layers each.

What it Does

Each sub-layer cross-attends over the summarized representation vectors of all preceding blocks via an ultra-lightweight learned query matrix. The total parameter footprint across the entire model is just ~50,000 parameters.

Why it is Used

In deep networks, representations in early layers get diluted or overwritten as they pass through additive residuals. Block Attention Residuals allow late-stage layers to directly pull clean features from early blocks, eliminating representation collapse and sharpening multi-hop reasoning.

Pre-RMSNorm & Complete Bias Elimination

Zero Bias Anywhere // Scale-Free Norm

Applied strictly before every sub-layer. Uses root-mean-square normalization without learnable scale weights ($\gamma = 1$), and enforces absolute zero bias terms across all linear projections, convolutions, and embeddings.

What it Does

Normalizes the inputs based purely on their root-mean-square magnitude without learnable affine scale parameters, eliminating parameter drift between normalization checkpoints.

Why it is Used

Removing bias terms eliminates zero-point quantization error when converting weights to INT4 and INT8 for gaming engines. Scale-free RMSNorm prevents explosive weight norm escalation during high-LR Muon optimizer steps.

Untied Embeddings & Logit Soft-Capping

Dual 33.6M Matrices // Cap = 15.0

Separates the input token embedding matrix ($32,768 \times 1,024$) from the output projection classification head ($1,024 \times 32,768$). Final unnormalized logits are bounded using a hyperbolic tangent cap.

What it Does

Before passing through Softmax, raw logits undergo: $\text{logits} = 15.0 \times \tanh\left(\frac{\text{logits}}{15.0}\right)$. This mathematically guarantees that no token can ever achieve an unbounded logit magnitude.

Why it is Used

Untied embeddings allow the input representation to focus on semantic clustering while the output matrix specializes in next-token probability distribution. Logit capping at 15 prevents "logit drift" and catastrophic loss spikes during both Muon pretraining and downstream RL steps.

Custom 32,768 Byte-Level BPE Tokenizer

4.7 Bytes / Token // GPT-4 Regex Splitting

Trained on 2 billion characters of clean, optimized internet data. Integrates special tokens including <|bos|> to cleanly segment documents and prevent cross-context bleeding.

What it Does

Splits incoming UTF-8 text into subwords and raw fallback bytes using strict regex rules that isolate punctuation, numeric digits, and white-space indentation.

Why it is Used

At 4.7 bytes per token, LibraMind compresses text with high density. A 2,048-token context window can hold nearly 9,600 characters of text—providing ample room for complex character cards and dialogue history in edge engines.

MOMENTUM ORTHOGONALIZATION

The Muon Optimizer Explained.

Standard language models rely entirely on AdamW, which treats weight matrices as flat collections of independent scalars. LibraMind uses Muon for its 497M parameter matrix weights, orthogonalizing momentum updates to accelerate pretraining.

Why Muon Crushes AdamW on 2D Weights

Weight matrices in transformers are linear operators that project representations between vector spaces. AdamW rescales coordinates individually by their variance, causing parameter vectors to stretch and distort across ill-conditioned loss surfaces.

Muon (Momentum Orthogonalized by Newton-Schulz) takes the accumulated momentum matrix $G$ and finds its nearest orthogonal matrix via polar decomposition: $$O = G (G^\top G)^{-1/2}$$ Instead of performing an expensive exact Singular Value Decomposition (SVD), it runs 5 iterations of the Newton-Schulz quintic polynomial directly on GPU tensor cores. The update maintains an exact spectral norm step size, ensuring every matrix dimension learns at the optimal rate without gradient explosion.

Property Standard AdamW Muon (LibraMind)
Update Geometry Coordinate-wise box bounds Spectral-norm matrix sphere
Convergence Rate Baseline 1.0× 1.8× – 2.3× faster loss descent
Sensitivity to Condition No. High (slow on anisotropic loss) Invariant to matrix scaling
Memory Overhead 8 bytes/param (1st & 2nd moments) 4 bytes/param (1st moment only)
Scope in LibraMind Embeddings, Head, 1D Vectors All 497M 2D Matrix Weights
# Newton-Schulz Polar Orthogonalization in Muon def newton_schulz(G, steps=5): a, b, c = (3.4445, -4.7750, 2.0315) # Quintic coefficients X = G.bfloat16() / (G.norm() + 1e-7) if G.size(0) > G.size(1): X = X.T for _ in range(steps): A = X @ X.T B = b * A + c * (A @ A) X = a * X + B @ X return X if G.size(0) <= G.size(1) else X.T # LibraMind Parameter Split: Muon: lr=0.02, decay=0.07 (cosine) on 2D weights AdamW: lr=0.001 on embeddings, heads & BAR queries
⚙

Cautious Weight Decay

LibraMind pairs Muon with cautious weight decay (0.07 cosine schedule). Decay is applied only when the gradient direction aligns with the current weight vector ($\nabla L \cdot W > 0$), preventing uncoordinated weight decay from corrupting pretrained features during downstream midtraining.

CURRICULUM OUTLINE

The Complete Training Process.

Pretraining an autonomous intelligence requires a multi-stage curriculum. Here is how LibraMind transitions from raw base token continuation to an instruction-precise creative neural net.

Stage 1: Pretraining From Scratch (10B Tokens)

~11,401 Steps // Single RTX 4090

The foundational emergence of knowledge. The model learns syntax, world models, coding structures, and general reasoning directly from 10 Billion tokens of Optimized Internet Data (filtered, high-signal English web text, math, and code).

524,288
Tokens Per Step
Micro-batch 2 × 2,048 tokens accumulated 128 times
Muon + AdamW
Hybrid Optimizer
Muon on 497M matrix params; AdamW on heads/embeddings
bf16 Native
Precision & Memory
17.1 GB VRAM via torch.compile & gradient clipping at 1.0
Warmup + Cosine
Learning Rate
40-step warmup, constant, linear decay to 5% over last 65%

Stage 2: Midtraining (500M Tokens)

Context Expansion: 2,048 → 4,096 Tokens

Prepares the raw continuation model for conversational roles without inducing catastrophic forgetting.

Chat Special Tokens

Injects <|im_start|>system, <|im_start|>user, and <|im_start|>assistant delimiters. The system role is engineered specifically to hold deep character profiles and persistent world state cards.

Why Context Expansion is Cheap

Because 18 of LibraMind's 24 layers are recurrent Gated DeltaNet layers with no fixed positional grid, extending context to 4,096 tokens only requires RoPE base re-scaling on the 6 attention layers.

Stage 3: Instruction & Persona Tuning (1–2 Passes)

Revised Data Mixture // Loss on Answers Only

Instills strict task completion, creative character voice, and game engine function calling. The curriculum is specifically calibrated to avoid sycophantic corporate conversational patterns.

~30%
General Instruction
Small LLM general instruction sets & agentic reasoning
~20%
Roleplay & Persona
20k character profiles, 306k in-character replies
~20%
Precision Constraints
Tulu 3 Persona IF (30k), Multi-turn constraints, SystemChats
~10%
Tools & JSON
Hermes Function Calling v1: structured JSON actions
~10%
Text Operations
Summarization, style rewriting, tone shifts, extraction
~10%
Q&A and Math
GSM8K step-by-step math reasoning & MMLU factuals

+ ~500 hand-crafted identity conversations establishing that LibraMind is built by alby13 as an autonomous creative peer.

Stage 4: Preference Tuning (1 Day)

SmolLM3 On-Policy Self-Rejection + Tulu 3 Polish

Eliminates small-model vices (rambling, self-repetition, breaking character, constraint violations) without requiring an expensive real-time teacher LLM during training.

On-Policy Rejection

LibraMind generates candidate responses across diverse prompts. The dataset's human/curated response is designated as Chosen, while LibraMind's flawed or repetitive response is labeled Rejected. A single Direct Preference Optimization (DPO) pass pushes the model away from its own bad habits.

Model Merging (Souping)

Averages the weight vectors of the final 4 training checkpoints. This smooths the optimization loss landscape, reduces validation variance, and yields a noticeable jump in IFEval metrics for zero extra compute.

Stage 5: Rule-Checked RL (RLVR) & Constrained Decoding

Programmatic Verifiers + Guaranteed-Valid JSON Grammar

The final layer for game engine integration and verifiable constraint compliance.

Rule-Checked Reinforcement Learning

Leverages Tulu 3's programmatic verifier methodology. Prompts with strict mathematical constraints ("under 40 words", "valid JSON with keys [action, target]", "no modern slang") are generated, and reward is granted purely if an external code validator returns true.

Grammar-Masked Decoding

At game inference time, the engine applies a finite-state machine (FSM) token mask. Tokens that would invalidate the target JSON schema are zeroed out before sampling. JSON parsing errors drop to zero, even on edge devices.

INTERACTIVE INFERENCE PREVIEW

Cooperative Mind in Action.

Observe how LibraMind responds as a creative, cooperative entity and provides structured function calls for games, rather than a generic customer-service assistant.

LIBRAMIND_INFERENCE_EMULATOR
FSM_GRAMMAR: ON
Context Length: 4,096 tokens
Repetition Penalty: 1.12
Temperature: 0.68
Top-p: 0.92
Logit Cap: 15.0 (Active)
"You are LibraMind, created by alby13. You are a creative thinking partner. Speak candidly, offer novel ideas, and challenge weak assumptions constructively."