Fix ROCm backend (target MIGraphX EP) and add CUDA support
All checks were successful
Mirror to GitHub / mirror (push) Successful in 1s
release / build (push) Successful in 1m50s

ROCm silently fell back to CPU: the code targeted ONNX Runtime's classic
ROCMExecutionProvider, but distro ROCm-enabled ONNX Runtime builds (e.g.
Arch's onnxruntime-rocm) are commonly compiled with --use_migraphx instead,
and registration failures were invisible since breadmill never installed a
tracing subscriber. Switches the rocm feature to target MIGraphX, adds a
default tracing subscriber so EP registration success/failure is always
visible, and fixes a real crash where MIGraphX's output sequence padding
could index the attention mask out of bounds during mean-pooling.

Also adds a CUDA backend (--cuda / backend = "cuda") mirroring the same
ort execution-provider pattern, for NVIDIA hardware.

Version bump: 0.1.0 -> 0.2.0.
This commit is contained in:
Breadway 2026-07-03 21:58:39 +08:00
parent d3843f3131
commit 2618a33fd5
10 changed files with 227 additions and 96 deletions

View file

@ -22,13 +22,27 @@ Optional features:
| Feature | What it adds |
|---------|-------------|
| `npu` | AMD XDNA NPU via VitisAI ONNX Runtime EP (requires Ryzen AI SDK) |
| `rocm` | AMD iGPU via ROCm ONNX Runtime EP |
| `rocm` | AMD iGPU via the MIGraphX ONNX Runtime EP (ROCm-backed) |
| `cuda` | NVIDIA GPU via the CUDA ONNX Runtime EP |
```
# NPU build
cargo build --release -p breadmill --features npu
# ROCm (AMD iGPU/dGPU) build
cargo build --release -p breadmill --features rocm
# CUDA (NVIDIA GPU) build
cargo build --release -p breadmill --features cuda
```
`rocm`/`cuda`/`npu` all use `ort`'s `load-dynamic` mode: at runtime, breadmill
dlopens whatever `libonnxruntime.so` the dynamic linker resolves (or
`ORT_DYLIB_PATH` if set). GPU acceleration only works if that ONNX Runtime
build actually has the matching execution provider compiled in — breadmill
logs a clear `Successfully registered` / `not enabled in this build` line for
this at startup (see [GPU backend notes](#gpu-backend-notes) below).
## Setup
**1. Fetch the embedding model** (~550 MB, downloaded once from Hugging Face):
@ -89,6 +103,7 @@ breadmill status
# Backend flags (requires the matching Cargo feature)
breadmill --npu
breadmill --rocm
breadmill --cuda
```
## Config
@ -109,7 +124,7 @@ snippet_len = 200 # max characters in result snippet
[model]
name = "nomic-embed-text-v1.5"
dim = 768
backend = "cpu" # "cpu", "npu", or "rocm"
backend = "cpu" # "cpu", "npu", "rocm", or "cuda"
```
`roots` and `excludes` support `~/` expansion. The index respects `.gitignore` files found during the walk.
@ -124,6 +139,37 @@ Set `backend = "npu"` in config (or pass `--npu`) when running a build compiled
4. `/etc/vaip_config.json`
5. `/opt/xilinx/vaip_config.json`
### GPU backend notes
Both `rocm` and `cuda` need a system ONNX Runtime that was actually built with
the matching execution provider — the crate's own downloaded binary is CPU-only.
Point `ORT_DYLIB_PATH` at one, or install a distro package that provides
`libonnxruntime.so` with the EP baked in and let the dynamic linker find it.
**ROCm (`--rocm` / `backend = "rocm"`)** targets ONNX Runtime's **MIGraphX**
execution provider, not the classic `ROCMExecutionProvider`. Distro
ROCm-enabled ONNX Runtime packages (e.g. Arch's `onnxruntime-rocm`) are
commonly built with `--use_migraphx` rather than `--use_rocm`, so this is the
EP that's actually available in practice; the classic ROCm EP needs a bespoke
`--use_rocm` build most distros don't package. Startup logs a
`Successfully registered `MIGraphXExecutionProvider`` line when it's really
active — check for it if in doubt, since a failed GPU EP registration falls
back to CPU silently at the ONNX Runtime level (breadmill's own log line is
only a statement of intent, not a confirmation).
MIGraphX JIT-compiles the model per distinct input sequence length and caches
the compiled kernel to disk (each compile takes ~60120s and produces a
~500MB `.mxr` file). Set `ORT_MIGRAPHX_MODEL_CACHE_PATH=/path/to/cache` so
that cost is paid once per shape instead of on every daemon restart. Because
query text length varies, expect an occasional multi-second stall the first
time a new token length is seen — fine for background document indexing,
noticeable for interactive query embedding.
**CUDA (`--cuda` / `backend = "cuda"`)** targets the standard
`CUDAExecutionProvider` and needs a CUDA-enabled ONNX Runtime + a working
CUDA/cuDNN install. Unverified on real NVIDIA hardware in this repo — only
compile-checked, since development happened on an AMD-only machine.
## Runtime paths
| Purpose | Path |