Package FreeToken (FlashML-org/FreeToken) — edge-native MoE serving
engine with OpenAI/Anthropic-compatible APIs, serving the `ft` CLI.
- torch-bin 2.11 (CUDA 12.9 wheel build) satisfies the torch>=2.11,<2.12
pin; extensions link the same cudaPackages.cuda_cudart via a synthetic
CUDA_HOME (cudart headers + nvcc crt/ headers, no nvcc needed).
- flashlib==0.3.0 vendored from the PyPI wheel (Triton-only; freetoken
never imports the CuTeDSL GEMM backends, so nvidia-cutlass-dsl is
omitted via pythonRemoveDeps).
- apache-tvm-ffi overridden to the pinned 0.1.13.post3 with a vendored
cython 3.3.0 (build needs >=3.2.8, nixpkgs has 3.2.4); pytest skipped
(upstream runs pytest-xdist with GPU tests).
- Unfree (CUDA EULA) is self-scoped: the package re-imports nixpkgs with
config.allowUnfree so no flake-level or user config change is needed.
- Optional accel extras (flashinfer/sglang-kernel) and the kernel-cache
wheel are not packaged; runtime falls back to pure-Triton kernels.
- Verified: nix build .#freetoken, ft --version, python imports check,
extension RPATHs.