# Tag

RDNA3

Posts with this tag: 12

← Todas las entradas
Measurement Gates that lie · part 2

Your Reference Fails Its Own Reference

A LoRA patch scored 89.06% on its worst prompt against a frozen 90% floor and looked FAILED. The unpatched base, run against the same floor for the first time, also failed at 89.06%, on a different prompt. The bar was unreachable by construction. Llama.cpp's own f16-KV configuration misses its own f32 reference the same way, at 5 of 7 tested lengths.

Read
GPU kernels A 35B MoE on one consumer card · part 3

Prefill as a Batch

A chunked prefill path for a 35B MoE passes identity on 23 of 23 prompts and runs 3.7 to 7.9 times faster than replaying the decode kernel one token at a time. Two variants that would have been faster still, WMMA attention and a dense SSM scan, reorder floating-point sums and stay off by default because they fail the same identity check.

Read
Systems design A 35B MoE on one consumer card · part 1

Experts in Host RAM

The routed experts of a 35B MoE moved to host RAM, VRAM footprint down to 2.68 GB from 21, output identical on 20 of 20 prompts. The first working version ran at 39 tok/s. Pinning the host store and reading misses in-kernel got it to 72.

Read
Measurement Gates that lie · part 1

A Gate the Candidate Can Write

An audit of this project's own optimization loop built four candidates designed to cheat its performance gate. One of them, a one-line deletion of a GPU synchronization call, passed at a fabricated plus sixty-three percent. The fix moved the stopwatch out of the file the candidate is allowed to touch.

Read