(1) inline stats on Attention/MLP/SSM nodes — '128h · GQA(8) · d=64 · win 4096' or 'h=4096 → 14336' rendered on the box itself, no need to expand children. (2) Backend auto-detects attention type (MHA / GQA / MQA / MLA via kv_lora_rank) and sliding_window / rope_scaling / head_dim into arch_features; Header info card shows 'attention: MLA · head_dim 128 · sliding window 4096'. (3) Hide '× ?' placeholders when repeat is non-numeric — looked broken; now nothing renders. Same for prototype 'N' fallback.
(1) detect residual `x = residual + ...` patterns in forward() and inject synthetic Add nodes between sibling modules — chain now reads norm → attn → ⊕ → norm → mlp → ⊕ instead of skipping the additions; (2) Add kind: small amber pill (140×40, rounded 14) like the canonical Add boxes in transformer paper diagrams; (3) Loop's `×N` tag becomes large pale-grey side text (font-size 64, weight 300) instead of a corner badge — closer to reference style. MoE keeps its compact 'K/N active' badge since that ratio is load-bearing.
conditional parent→all-kids: only for parallel-tower parents (2+ tower attrs like visual+language_model). Sequential parents (BERT bert→embeddings→encoder→pooler) get parent→first only — fixes BERT regression where embeddings/encoder/pooler were collapsing onto the same rank side-by-side instead of stacking vertically.
Loop/Model containers: connect parent to ALL ordered kids in dagre (not just first) — fixes parallel towers in Qwen3.5-MoE-VL where language_model had no constraint to its parent and floated to rank 0 overlapping INPUT_TERM. Now both visual and language_model are pinned to parent.rank+1, dagre places them side-by-side at the same rank.
explicit horizontal separation for parallel towers — after dagre layout, if visual/language_model bboxes overlap (or sit too close), translate the right one to clear the overlap with an 80px gap; descendants and nested loop boxes move with it. Fixes Qwen3.5-MoE-VL where vision and text towers were drawn on top of each other in the middle of the canvas
(1) tighten MoE detection — suffix-based match (MoE/MoeBlock/SparseMoeBlock/Experts) + Expert-not-Vision substring; fixes Qwen3_5MoeVisionBlock and Qwen3_5MoeDecoderLayer being false-positive MoE just because the model family is named *Moe*. (2) extend SSM kind to GatedDeltaNet / *LinearAttention / *TimeMix — Qwen3.5-MoE's linear_attn is now teal Selective SSM, not unclassified Other
multimodal towers as parallel containers — visual / vision_tower / audio_encoder / language_model / text_model attrs each get their own model_container box, parent's bbox excludes them so they sit alongside, and no sibling-chain edges link two towers (they're parallel paths until a merger). Fixes Qwen3.5-MoE-VL where vision and language model were force-chained vertically inside one big container.
forward order recurses through helper methods — CLIPModel.forward calls get_image_features/get_text_features which then call vision_model/text_model/visual_projection/text_projection. Walker now recurses into any self.X() where X is also a method of the same class, accumulating module calls in source order. Same fix applied to runtime variant.
surface model-wide FLOPs/token: /api/load and /api/compare now include `flops_per_token` from the root graph node; OutputPanel shows '≈ 2.5 GFLOPs / token' callout pill at the bottom; Header model-info card adds a 'compute' line
forward-order walks ALL forward-named methods of a class — Mamba2Mixer.forward only dispatches to torch_forward/cuda_kernels_forward; walking just `forward` returned [] so conv1d/in_proj/out_proj had no forward edges and fanned out horizontally. Now we collect self.X calls from forward + cuda_kernels_forward + torch_forward + slow_forward + fast_forward (any *forward* method), skipping dispatching self.<method>_forward calls
dedupe `self.X = ...` reassignment across both `if` branches — when DeepSeek/GLM has `if first_k_dense_replace: self.mlp = DenseMLP else: self.mlp = MoE`, both branches were emitting and creating two nodes with the same path. Now: keep the higher-priority branch (MoE/Block > Attention > MLP/Norm > Linear) so the interesting architecture wins.
MoE detection for fused experts wrappers (DeepseekV4Experts, GroupedLinear, etc.) — any attr named `experts` whose class is MLP/Other gets reclassified as MoE with router info from config; *Router / *Gate suffix classes also get the 🎯 router badge
MoE router annotation: (1) extract top_k/num_experts/scoring_func/norm_topk_prob from config and attach to MoE container's args; (2) tag `gate`/`router` Linear attrs with is_router so they get a fuchsia '🎯 MoE ROUTER' badge; (3) MoE container shows 'router · softmax → top-8' info inline
add Transformer vs Mamba quick-compare + STATE-SPACE (SSM) preset category — Mamba-Codestral-7B and mamba-2.8b-hf accessible in one click; comparison loads Qwen2.5-1.5B (transformer) on the left and Mamba-Codestral on the right to show selective-SSM blocks vs attention
lift tree edges through invisible ancestors — when a kind=Other wrapper (Mamba2Mixer etc.) is filtered, its visible descendants attach to the nearest visible ancestor instead of orphaning
add 'Mixer' to MLP patterns — Mamba2Mixer (the SSM core) is now MLP kind, so its conv1d/in_proj/out_proj children no longer orphan when the parent class is unclassified
(1) FLOPs popup on hover — every node now has a native tooltip with role, class, path, params, FLOPs/token, matmul shape; plus a small ≈XGF badge in the top-left corner of non-narrow nodes (2) GitHub link in NodeCard source section — '📎 github' button next to the source viewer opens the modeling .py file at the exact line where the class is defined
(1) MoE/Loop prototype gets a '✱ 1 of N' / '🔁 1 of N' badge — makes the single example expert/layer obvious inside its container; (2) T5LayerFF / DenseActDense / DenseGatedActDense classified as MLP not Block (matched 'Layer' before); (3) scan_kinds.py utility to audit class→kind mappings across architectures
(1) classify Head only by SUFFIX (BertLMPredictionHead = Head, BertPredictionHeadTransform = Other) so transform sub-blocks don't get pulled into the head ; (2) move Head to ATOMIC_KINDS — head wrappers no longer expand into nested forest of cls→predictions→transform→decoder; one clean Output Projection box at the end. Click → NodeCard subtree shows internals on demand
swap Llama-3.2-1B for Qwen2.5-1.5B-Instruct in presets — Dense vs MoE now compares same-family Qwen2.5 vs Qwen3-MoE; 'Llama vs Qwen' renamed to 'Phi-2 vs Qwen2.5'
categorized presets (Dense LLM / MoE / Encoder / Audio / Vision-MM / Encoder-Decoder) + Quick Compare grid (Dense vs MoE, BERT vs ModernBERT, Llama vs Qwen, Whisper tiny vs turbo, ViT vs CLIP, T5 vs BART) — one click loads both halves
ModernBERT-style heads: (1) `decoder` Linear at depth 1 reclassified as Head (only when the class is Linear, so BART decoder Module is unaffected); (2) isOutputAttr also accepts kind=Head so any head-classified node goes to the column bottom; (3) when multiple heads exist (head + decoder), stack them vertically beneath the model container instead of overlapping
model comparison: /api/compare endpoint + ComparePanel — paste a second model id, see side-by-side diff of params, layers, hidden, vocab, head_kind, MoE expert counts; differing rows highlighted with amber border
NeMo composite detection (NVIDIA Canary etc.) — config has `perception` block with non-transformers audio encoder; we now (1) override modality to audio, (2) extract sample_rate/n_mels from perception.preprocessor for InputPanel, (3) show a yellow ⚠️ note in the header explaining only the LLM portion is visualized
forward-order parsing now sorts by source line — fixes Whisper showing 'decoder → encoder' (encoder lived inside an `if` block; ast.walk's BFS-by-structure put it after the top-level decoder call). Applied to both static and runtime forward-order helpers
fix black screen: move all NodeCard hooks (useState/useEffect) above the conditional `return null` — React was throwing 'rendered more hooks than during the previous render' on the first model load, blanking the canvas
(1) FLOPs/token estimate per node — Linear=2·in·out, Conv=2·Cin·Cout·K², Norm=5·H, Activation=H, with Loop×N and MoE×K_active aggregation up the tree; shown in NodeCard pill (2) /api/source endpoint + 'show source' button in NodeCard — fetches the original class source from the loaded model's modeling file and renders in a code block
(1) descendants() walks through invisible intermediate wrappers — Mistral3 vision_tower.transformer (kind=Other, hidden) was blocking the model-container bbox from reaching the actual transformer Block subtree. (2) drop 'Projector' from Head patterns — multi_modal_projector is an MLP-like bridge, not an output head; output heads are detected by attr name only