Allow a pinned JIT SDK without rebuilding PyTorch or the extension.
Keep runtime device bitcode unchanged: exporting TheRock bitcode into
the serving process crashes CLR 7.2 on a basic tensor operation.
Disable unsupported expandable segments and check all five HIP kernel
modules, with an optional GPU matmul and module-loading regression test.
env.sh assumes a distro box: a system cc for Triton's runtime C builds, a hipcc
for the JIT HIP kernels and the repo root on PYTHONPATH. The wrappers now export
all three (, , for the amdgcn bitcode,
PYTHONPATH=${src}/gr for gr_mix_hip).
The wheel ships only the files pyproject lists, so the .hip kernel sources, the
qsa_proof marker gating the QSA prefill kernel, the dense-GEMM tuning seed and
gr/ never reached site-packages: kernels fell back to Triton and every start
re-tuned dense GEMM for 7 minutes.
Verified on the Strix Halo box: server READY, start-up 713 s -> 103 s, warm-up
418 s -> 4 s, 584-token prompt served over /v1/chat/completions with no
'unavailable' fallbacks.
Kyojin (Yamz-Labs/kyojin) is a fork of ExLlamaV3 that runs 100 GB-class
EXL3 MoE packs on one 128 GB Ryzen AI Max box. It ships a torch C++
extension that torch's cpp_extension cross-compiles to HIP, and upstream
supports gfx1151 only, so the kernels are built for gfx1151 alone and
the package is marked Linux-only. Pinned to the v1.3 release.
The extension must be compiled against a ROCm build of torch. The flake's
nixos-unstable carries torch 2.13, whose ROCm build fails in nixpkgs (the
CK SDPA configure step runs a script through /bin/bash) and has no binary
on the cache, so the build would compile PyTorch from source and die.
python3Packages.torchWithRocm from the already pinned nixpkgs-torch211
input gives torch 2.11 with ROCm 7.2.2, gfx1151 in its target list and a
cached binary, so default.nix builds through that input the way freetoken
does.
nixpkgs splits the single devel tree the AMD wheels ship, so two
symlinkJoins stand in for it: one as ROCM_HOME (hipcc, headers, amdgcn
bitcode) and one handed to setup.py as EXL3_ROCM_DEV_INCLUDE, which wants
hipsparse/, rocsparse/, rocrand/, thrust/ and pybind11 (nixpkgs torch no
longer exports pybind11 headers). hipsolver is in there because torch's
own HIPContextLight.h includes it.
setup.py gets one patch in postPatch. It imports exllamav3 to reach
build_config, that runs the package __init__, which imports ext.py, which
JIT-compiles the entire extension whenever no precompiled exllamav3_ext is
importable, meaning every wheel build, into a $HOME the sandbox does not
grant. Registering a stub package keeps the two leaf modules importable
without that detour.
doCheck = false: tests/ drives a real gfx1151 GPU and the published
packs. Inference on a card is not verified here, this builder has no
/dev/kfd, so the wheel-only LD_PRELOAD workaround for torch's bundled HSA
runtime is untested. It should not be needed, the store torch links the
store rocm-runtime.
Verified with nix build .#kyojin at v1.3: exllamav3_ext loads, torch
reports hip 7.2.53211, the closure holds one torch (ROCm), and
kyojin-serve-qwen plus kyojin-serve-glm print their usage.
Refs nix-overlay-du9