Commit Graph

1 Commits

Author SHA1 Message Date
ab88b09ceb feat(packages): add freetoken 0.1.2 NVIDIA MoE inference runtime
Some checks failed
CI / check (push) Has been cancelled
Package FreeToken (FlashML-org/FreeToken) — edge-native MoE serving
engine with OpenAI/Anthropic-compatible APIs, serving the `ft` CLI.

- torch-bin 2.11 (CUDA 12.9 wheel build) satisfies the torch>=2.11,<2.12
  pin; extensions link the same cudaPackages.cuda_cudart via a synthetic
  CUDA_HOME (cudart headers + nvcc crt/ headers, no nvcc needed).
- flashlib==0.3.0 vendored from the PyPI wheel (Triton-only; freetoken
  never imports the CuTeDSL GEMM backends, so nvidia-cutlass-dsl is
  omitted via pythonRemoveDeps).
- apache-tvm-ffi overridden to the pinned 0.1.13.post3 with a vendored
  cython 3.3.0 (build needs >=3.2.8, nixpkgs has 3.2.4); pytest skipped
  (upstream runs pytest-xdist with GPU tests).
- Unfree (CUDA EULA) is self-scoped: the package re-imports nixpkgs with
  config.allowUnfree so no flake-level or user config change is needed.
- Optional accel extras (flashinfer/sglang-kernel) and the kernel-cache
  wheel are not packaged; runtime falls back to pure-Triton kernels.
- Verified: nix build .#freetoken, ft --version, python imports check,
  extension RPATHs.
2026-08-30 20:51:15 +03:00