RefineFuture AI
RefineFuture AI

Language models, compiled to the metal.

We develop and provide a new approach of running LLM/LMs’ inference/training on GPU/NPU backends through C++ implementation and compile for High-Performance, Flexibility and Easy-to-Use

refft-hexagon — Dragonwing IQ-9075 (QCS9075) 1×HTP
883.4tok/s
Peak prefill speed
99.8tok/s
Peak decode speed
145ms
Best TTFT

The terminal replays the measured who_are_you run at its measured decode rate. The figures above are the peak per axis and come from different prompts by nature. Peak figures for Dragonwing IQ-9075 (QCS9075) 1×HTP, read live from perf.yaml.

What we build

One runtime, compiled per backend. No Python in the serving path, no framework to install on the device, and every published number reproducible on hardware you can buy.

Pure C/C++, no dependencies

Inference, training and serving in one toolkit, built to run with minimal setup on a wide range of hardware — locally and in the cloud.

GPU and NPU backends

Custom kernels per target: CUDA for NVIDIA, HIP for AMD, MUSA for Moore Threads, and Hexagon for Qualcomm NPUs — where the whole pipeline runs on the accelerator.

Measured, not estimated

Every model ships as a self-contained package with a benchmark anyone can re-run. A platform that does not work is recorded as a result, not omitted.

Run a model on device

Install the runtime and send a prompt. Two lines, no build step.

Install and run

Full instructions, serving and every supported platform live in the runtime’s README.

RefineFuture AI