Open work / 04
Built to be useful.
The public benchmark notebook, upstream contributions, and a carefully bounded compatibility investigation.
Historical measurementsThe primary public collection for LLM and workload measurements, methodology, and corrections. The website curates a small reviewed subset; the repository's older prose may retain unresolved claims.
Browse reviewed records →Upstream workSix changes corroborated in retained upstream history, covering Q8_0 reorder, repeated-prompt correctness, allocation, alignment, BF16 decode, and subgroup sizing.
Read the case study →ExperimentalB70-CUDA investigation
A source-first, hybrid approach: translate suitable CUDA kernels and use native XPU operations where they fit. The retained Mamba-130M alpha checkpoint is narrow, depends on a development tree and bridge, and does not establish general Mamba or CUDA support.
There is no verified public B70-CUDA release or repository link in this evidence collection. Use the practical lessons as a porting checklist.
Explore the decision guide →CONTRIBUTING EVIDENCEMake the next result reproducible.
A useful contribution includes the model revision, build, environment, exact workload, repeated samples, and output validation. Report a failed configuration as carefully as a fast one.
Submit a result or correction ↗