Tensorunfolded.

unfold.py
1def unfold(T: nuke.Tensor, mode: fp16) -> nuke.Tensor:
2    """U_Ω^(n)(T) = ∮ exp(i/ħ ∫ A_μ dx^μ) ⋆ (∇_ξ T ∧ J)"""

Neural Unfold Kernel Engine

NukeTorch starts from a different idea of the tensor, a new mathematical formulation we will publish, and builds its kernels from it. Same models, same weights, same output, less in the way of the hardware - the fastest AI inference engine. We started with Apple silicon, and we are measuring every step in public.

Voice in, voice out. Both faster on a Mac.

NukeTorch runs open text-to-speech and speech-to-text models on Apple silicon faster than MLX, the fastest alternative on a Mac today. Same model files, no changes to your code, and output that passes the same quality checks.

Text to speech · Kokoro, F5-TTS, Fish Speech
about 3×
Kokoro against MLX · 2.5× per request, 3.4× in bulk
  • F5-TTS takes a third less time per clip
  • Fish Speech: 1.2× per clip and 4.4× in bulk, at the same precision as MLX
  • Kokoro and F5 need about a quarter of the memory; on a 1,000-sentence F5 job, 1.7 GB against MLX’s 106 GB
Speech to text · Parakeet, Whisper
about 3×
Parakeet against MLX · 2.3× per clip, 3.4× on 1,000 clips
  • Whisper Tiny 2.5× faster on 1,000 clips
  • Whisper large-v3-turbo 1.8× faster on 1,000 clips, at the same precision as MLX
  • Parakeet needs 95% less memory on a bulk job
  • The same word error rate on 1,000 real recordings

Measured on one Apple M4 Max with the same model files the authors publish. No NVIDIA results yet. Each product page lists what we have not measured.

How it works

Your model is waiting on its dispatcher

Every operation a model performs has to be routed: looked up, allocated for, launched. PyTorch's dispatcher does that job well for the case it was designed around, which is training large models on datacenter hardware, where the routing cost disappears next to the arithmetic.

Speech is the opposite case. A second of audio, generated or transcribed, is thousands of tiny operations, each a few microseconds of real arithmetic that has to be routed before it can run. On small models, the hardware spends much of its time waiting to be told what to do next. NukeTorch replaces that routing layer. The model files, the arithmetic and the output stay the same.

PyTorch, eager where most readmes send Mac users
MLX most of the gap already closed, and the bar we measure against
NukeTorch same arithmetic, little between it and the hardware
Schematic, not measured data. Coloured blocks are arithmetic; the field between them is hardware waiting to be told what to run next. We have not yet published a measurement that isolates that share, which is why this figure is drawn rather than plotted. Every number on this page is measured end to end.

Anticipating

Questions we expect

Including the ones that do not flatter us. What each set of results does and does not show is on the text to speech and speech to text pages.

What is this not good for?

Inference, not training. The models we have ported, not any model you bring. Apple silicon, because the current implementations are C++ and Metal. Anything outside the supported path falls back to an MLX worker rather than failing.

It is also not a framework. You cannot write a new model in it tomorrow, and that is deliberate rather than unfinished.

Is this just MLX with extra steps?

It is the right question, and it is why MLX is the comparison in every headline here. NukeTorch does not sit on top of MLX. On the same machine, Kokoro runs about 3× faster (2.5× on one request, 3.4× in bulk, and still 2.35× at MLX’s own full-precision settings), F5-TTS takes a third less time per clip with identical sampler settings, and Fish S2 Pro, at the same precision on both sides, is 1.21× faster per clip. Parakeet transcribes 2.3× faster per clip and 3.4× faster across 1,000 recordings, and Whisper large-v3-turbo, again at identical precision, is 1.8× faster across 1,000. If the answer were "use MLX", these tables would say so.

Why not contribute this upstream to MLX or PyTorch?

Because the thing that costs time is the thing those projects exist to provide. A dispatcher that can route any operation to any backend for any data type, and that you can write new models against tomorrow, is valuable and not free. You cannot remove the per-operation cost without removing the generality that makes it worth using.

What we built is narrow on purpose: a fixed set of models, one hardware family, inference only. It is a different trade, and it only pays on workloads shaped like this one.

How do I check any of this myself?

Every model has a benchmark report with the pinned checkpoints, the environment, the procedure and the limits of each number: Kokoro, F5-TTS, Parakeet, Whisper large-v3-turbo and Whisper Tiny.

How is it licensed?

Undetermined.

What about NVIDIA?

We have nothing to show you yet. It is the next thing we build, and until it exists we are not going to make claims about datacenter serving, cost per minute, or anyone's margin.

Do I have to change my model code?

No. Same model files, same checkpoint revision, same sampler defaults. Anything outside the supported path falls back rather than failing: for F5, text outside English or a reference that is not mono PCM16 goes to an MLX worker instead.

Why will you not publish the internals?

It is a commercial decision rather than a technical one. What we can do instead is make everything around the claim checkable: pinned revisions, published scripts, receipts, output gates fixed before the build existed, and known-fail controls that prove the gates bite.

Get in touch

For a commercial demo or a pilot, write to us and say which model you care about.