feat(kyojin): add Kyojin ExLlamaV3 engine for AMD Strix Halo
CI / check (push) Has been cancelled

Kyojin (Yamz-Labs/kyojin) is a fork of ExLlamaV3 that runs 100 GB-class
EXL3 MoE packs on one 128 GB Ryzen AI Max box. It ships a torch C++
extension that torch's cpp_extension cross-compiles to HIP, and upstream
supports gfx1151 only, so the kernels are built for gfx1151 alone and
the package is marked Linux-only. Pinned to the v1.3 release.

The extension must be compiled against a ROCm build of torch. The flake's
nixos-unstable carries torch 2.13, whose ROCm build fails in nixpkgs (the
CK SDPA configure step runs a script through /bin/bash) and has no binary
on the cache, so the build would compile PyTorch from source and die.
python3Packages.torchWithRocm from the already pinned nixpkgs-torch211
input gives torch 2.11 with ROCm 7.2.2, gfx1151 in its target list and a
cached binary, so default.nix builds through that input the way freetoken
does.

nixpkgs splits the single devel tree the AMD wheels ship, so two
symlinkJoins stand in for it: one as ROCM_HOME (hipcc, headers, amdgcn
bitcode) and one handed to setup.py as EXL3_ROCM_DEV_INCLUDE, which wants
hipsparse/, rocsparse/, rocrand/, thrust/ and pybind11 (nixpkgs torch no
longer exports pybind11 headers). hipsolver is in there because torch's
own HIPContextLight.h includes it.

setup.py gets one patch in postPatch. It imports exllamav3 to reach
build_config, that runs the package __init__, which imports ext.py, which
JIT-compiles the entire extension whenever no precompiled exllamav3_ext is
importable, meaning every wheel build, into a $HOME the sandbox does not
grant. Registering a stub package keeps the two leaf modules importable
without that detour.

doCheck = false: tests/ drives a real gfx1151 GPU and the published
packs. Inference on a card is not verified here, this builder has no
/dev/kfd, so the wheel-only LD_PRELOAD workaround for torch's bundled HSA
runtime is untested. It should not be needed, the store torch links the
store rocm-runtime.

Verified with nix build .#kyojin at v1.3: exllamav3_ext loads, torch
reports hip 7.2.53211, the closure holds one torch (ROCm), and
kyojin-serve-qwen plus kyojin-serve-glm print their usage.

Refs nix-overlay-du9
This commit is contained in:
2026-10-07 16:14:09 +03:00
parent ff1fecc0ed
commit e061d3c170
3 changed files with 218 additions and 0 deletions
+2
View File
@@ -31,6 +31,7 @@ A custom Nix overlay and flake providing additional packages not found in upstre
| `shardr` | Decentralized LLM repository: content-addressed model store, BitTorrent-based sync, OpenAI-compatible serving via llama-server | AI Inference |
| `skillsmcp` | MCP server that exposes Agent Skills to AI agents via the Model Context Protocol | MCP Servers |
| `kubernetes-mcp-server` | Model Context Protocol (MCP) server for Kubernetes and OpenShift | MCP Servers |
| `kyojin` | ExLlamaV3-based inference engine with speculative decoding for 100 GB-class MoE packs on one 128 GB AMD Strix Halo (gfx1151, ROCm 7) | AI Inference |
| `llama-cpp` | llama.cpp pinned to build b11439, built for Strix Halo (gfx1151) with ROCm + Vulkan backends | AI Inference |
| `loop` | Corporate messenger for your team | Communication |
| `radar` | Modern Kubernetes visibility — topology, event timeline, service traffic, resource browsing, Helm management, and GitOps support | Kubernetes |
@@ -199,6 +200,7 @@ nix-overlay/
│ ├── haivemind/ # Multi-model AI consensus runner
│ ├── hipengine/ # ROCm-native LLM inference engine for AMD GPUs
│ ├── kubernetes-mcp-server/ # MCP server for Kubernetes and OpenShift
│ ├── kyojin/ # ExLlamaV3-based inference engine for Strix Halo (gfx1151)
│ ├── loop/ # Corporate messenger for your team
│ ├── mcp-gateway/ # MCP protocol gateway
│ ├── omniroute/ # Unified AI router with 160+ providers