Tensorunfolded.

unfold.py
1def unfold(T: nuke.Tensor, mode: fp16) -> nuke.Tensor:
2    """U_Ω^(n)(T) = ∮ exp(i/ħ ∫ A_μ dx^μ) ⋆ (∇_ξ T ∧ J)"""

Neural Unfold Kernel Engine

NukeTorch starts from a different idea of the tensor, a new mathematical formulation we will publish, and builds its kernels from it. Same models, same weights, same output, less in the way of the hardware - the fastest AI inference engine. We started with Apple silicon, and we are measuring every step in public.

Voice in, voice out. Both faster on a Mac.

NukeTorch runs open text-to-speech and speech-to-text models on Apple silicon faster than MLX, the fastest alternative on a Mac today. Same model files, no changes to your code, and output that passes the same quality checks.

Text to speech · Kokoro, F5-TTS, Fish Speech
about 3×
Kokoro against MLX · 2.5× per request, 3.4× in bulk
  • F5-TTS takes a third less time per clip
  • Fish Speech: 1.2× per clip and 4.4× in bulk, at the same precision as MLX
  • Kokoro and F5 need about a quarter of the memory; on a 1,000-sentence F5 job, 1.7 GB against MLX’s 106 GB
Speech to text · Parakeet, Whisper
about 3×
Parakeet against MLX · 2.3× per clip, 3.4× on 1,000 clips
  • Whisper Tiny 2.5× faster on 1,000 clips
  • Whisper large-v3-turbo 1.8× faster on 1,000 clips, at the same precision as MLX
  • Parakeet needs 95% less memory on a bulk job
  • The same word error rate on 1,000 real recordings

Measured on one Apple M4 Max with the same model files the authors publish. No NVIDIA results yet. Each product page lists what we have not measured.

How it works

Your model is waiting on its dispatcher

Every operation a model performs has to be routed: looked up, allocated for, launched. PyTorch's dispatcher does that job well for the case it was designed around, which is training large models on datacenter hardware, where the routing cost disappears next to the arithmetic.

Speech is the opposite case. A second of audio, generated or transcribed, is thousands of tiny operations, each a few microseconds of real arithmetic that has to be routed before it can run. On small models, the hardware spends much of its time waiting to be told what to do next. NukeTorch replaces that routing layer. The model files, the arithmetic and the output stay the same.

PyTorch, eager where most readmes send Mac users
MLX most of the gap already closed, and the bar we measure against
NukeTorch same arithmetic, little between it and the hardware
Schematic, not measured data. Coloured blocks are arithmetic; the field between them is hardware waiting to be told what to run next. We have not yet published a measurement that isolates that share, which is why this figure is drawn rather than plotted. Every number on this page is measured end to end.

Anticipating

Questions we expect

Including the ones that do not flatter us. What each set of results does and does not show is on the text to speech and speech to text pages.

What is this not good for?

We're expanding to more models, but we're a small team with limited capacity, and right now every model has to be ported by us. As the team grows, we plan to make the framework much easier for others to use.

Is this just MLX with extra steps?

No. NukeTorch is a new engine, built from the ground up on our own mathematical principles. We pursue one question: what is a number? We used as few open-source libraries as we could; we don't even use NumPy in the engine, though some of our benchmark scripts and weight converters do. We learned from others, of course, so we won't claim every idea in our codebase is new. But we wrote every line ourselves, our intention was clear, and we exceeded our goal.

Why not contribute this upstream to MLX or PyTorch?

We'd love to, and we'd love to work closely with both teams. But for today, let us enjoy being the fastest engine on Apple silicon for every model we've measured.

How do I check any of this myself?

Every model has a benchmark report with the pinned checkpoints, the environment, the procedure, and the limits of each number. For now, we run private demos on request.

How is it licensed?

We're working on it. For now, it's undetermined.

What about NVIDIA, AMD, or TPU?

We'd love to work with them. Let's just say we have plans.

Do I have to change my model code?

Nope. But if you want every last bit of performance, we can work with you.

Will you publish the internals?

We'll share our position on this soon.

Kokoro runs about 3× faster. F5 takes a third less time. Both need about a quarter of the memory.

The fastest way to run these models on a Mac today. Same model files, no changes to your code, and output that passes the same quality checks.

Kokoro-82M · one request and in bulk
One request
18.5 ms
vs 46.7 ms · 2.5×
1,000 sentences
46 s
vs 156 s · 3.4×

Against MLX. A single sentence, and a bulk job of 1,000 real sentences.

F5-TTS v1 Base · time per clip
1.37 s
vs 2.05 s with MLX · 33% less

Identical sampler settings on both sides. In bulk, 44% less time across five passes.

Peak memory footprint
Kokoro
0.54 GB
vs 2.05 GB
F5-TTS
0.88 GB
vs 3.22 GB

Peak physical footprint on a single request, NukeTorch against MLX, measured the same way for both.

Measured on one Apple M4 Max with the same model files the authors publish. Single requests: medians of 60 runs on a short sentence. Bulk: 1,000 real sentences from the LibriSpeech corpus. Memory measured with the same outside probe for every engine. Output passes the same automated quality checks, and our own ears; no formal listening study yet. No NVIDIA results. What we have not measured. Why it is faster: how NukeTorch works.

Why it matters

The difference

The same three results land differently depending on what you build. These are the people we think they matter to most.

2.05 GB → 0.54 GB

You ship a Mac app with a local voice

A quarter of the memory lowers the machine your app needs, so more of your users' Macs qualify. The audio never leaves the device, and there is no per-minute cloud bill.

Kokoro, peak memory footprint
47 ms → 18.5 ms

You run a local voice assistant

Every step between a person finishing a sentence and hearing a reply adds its own wait: transcription, the language model, then speech. The voice starts sooner, and the memory it gives back goes to the language model, which is usually what runs out first.

Kokoro, time to first audio
60 min → ~33 min

You generate speech in bulk

Voice cloning, narration and dubbing with F5: in our bulk test of 1,000 real sentences, work that takes an hour under MLX took about 33 minutes. With Kokoro, the same 1,000 sentences take 46 seconds instead of 2.6 minutes. Passages of several minutes are next on our test list.

F5-TTS and Kokoro, 1,000-sentence bulk job

Evidence, in one table

Three models, one Mac

MLX is the comparison we hold ourselves to, because it is the fastest alternative on a Mac. PyTorch appears in each model's table because it is the reference the model authors publish against.

Model One request vs MLX 1,000 sentences vs MLX Memory vs MLX Evidence
Kokoro-82M 2.5× faster 3.4× faster 74% less Mixed precision
F5-TTS v1 Base 33% less time 44% less time 73% less Mixed precision
Fish Speech S2 Pro 1.2× faster 4.4× faster 18% less Same precision
Time to first audio

How long the listener waits before the voice starts. In these tests the whole clip returns at once, so it is also the time to the finished clip.

Bulk job

1,000 different sentences from the LibriSpeech corpus, generated back to back. NukeTorch runs in its default bulk mode, which schedules the requests itself; MLX generates them one after another, as it ships.

Real-time factor

Seconds of audio produced per second of computing. Above 1.0, audio is generated faster than it plays.

Peak memory footprint

The most memory the model used, as Activity Monitor reports it. It decides how many copies fit on a machine, and what else can run alongside.

Text to speech · model 1 of 3

Kokoro-82M

About 3× faster than MLX: 2.5× on a single request and 3.4× on a bulk job of 1,000 sentences. The model needs about a quarter of the memory, and it stays well ahead even when run at MLX’s own full-precision settings.

Mixed precision

Maintained by hexgrad · Apache 2.0 · v1.0, voice af_heart, seed 0 · One request: 45,600 samples at 24 kHz (about 1.9 seconds), medians of 60 runs · Bulk: 1,000 LibriSpeech sentences, median of 5 passes

MetricNukeTorchMLXPyTorchvs MLXvs PyTorch
One request, text to finished audio (ms)18.4746.6561.682.53×3.34×
1,000 sentences in bulk (s) §46.36156.38not run3.37×n/a
Real-time factor, one request (×)102.8540.7330.802.53×3.34×
Real-time factor, bulk (×)150.4244.57not run3.37×n/a
Transcription error rate, bulk †2.25%2.12%not run0.13 points moren/a
Peak memory, one request (MB)5362,0482,15074% less75% less
Peak memory, 1,000 sentences (MB) §2,3557,987not run71% lessn/a

§ NukeTorch in its default bulk mode, which schedules the requests itself; MLX generating them one after another, as it ships. Median of five passes; the slowest NukeTorch pass was 46.38 s and the fastest MLX pass 156.27 s. PyTorch was not run in bulk. Memory is the process’s physical footprint sampled from outside, the same probe for every engine, in runs separate from the timings; MLX keeps memory it has finished with in its own cache, so its bulk peak reflects its largest working set. † The difference comes from NukeTorch’s own text-to-phoneme step, not the synthesis: fed the same Misaki phonemes as MLX, NukeTorch scores 2.14% against 2.12%.

What happens technically

Kokoro is small enough that very little of a request is arithmetic. On the test clip a single request finishes after 18.5 ms instead of 46.7 ms under MLX. The work being done is identical; what is gone is the routing around each operation. In bulk the gap widens to 3.4×, because NukeTorch also schedules the requests itself while MLX works through them one at a time. The five bulk passes landed within a third of a second of each other, 46.1 to 46.4 seconds, so this is a steady result rather than a lucky run. The memory result has the same cause. A general framework keeps a graph, an allocator and its own intermediate buffers alive for every operation it might be asked to run next, and none of that is needed when the sequence of operations is fixed in advance.

What it means for you

On a single sentence on a fast Mac, the time saved is small in absolute terms. It adds up when Kokoro is one step in a chain. In a local voice assistant, transcription, the language model and speech each add a wait, and about 28 ms back from the last step is 28 ms nobody sits through. For bulk narration, 1,000 sentences take 46 seconds instead of 2.6 minutes.

The memory matters in every setup. About 0.54 GB instead of 2 GB is room for a larger language model on the same Mac, or a lower minimum spec for an app that ships Kokoro. If you maintain Kokoro, it is also a better answer for your readme's Mac section than PyTorch's Apple fallback.

Does the output still match

Qualified: 42 requests across four languages. Waveform cosine similarity of at least 0.999998 against the reference, where the gate was set at 0.999. Relative RMSE at most 0.20%, against a 5% gate. Identical sample counts. Word error rate from automatic transcription equal on all 36 English requests. Deliberately broken control cases fail the same gates, so the gates are doing work.

On the 1,000-sentence bulk test, NukeTorch’s output has 2.25% of words wrong against MLX’s 2.12%, 0.13 percentage points more. The cause is NukeTorch’s own text-to-phoneme step, which follows Misaki’s rules without its part-of-speech tagger and without its fallback for words outside the dictionary. Given the same Misaki phonemes, NukeTorch scores 2.14%, level with MLX.

Same settings as MLX · Kokoro at full precision

To answer the obvious objection, that we beat slow settings, NukeTorch was also run exactly the way MLX and PyTorch run: FP32 weights, the same Misaki text-to-phoneme step, and one request at a time in bulk. It is still 2.35× faster than MLX on a single request (19.88 ms against 46.65 ms), 3.10× faster than PyTorch, and 2.58× faster than MLX across the 1,000 sentences (60.6 s against 156.4 s). Transcription error rates match: 2.14% against 2.12%. Peak memory on a single request is 636 MB against MLX’s 2,048 MB. Most of the speed is the runtime itself; half precision and NukeTorch’s own bulk scheduling add the rest.

Nuance
  • The single-request clip is short: about 1.9 seconds of audio from a three-word sentence. Short clips flatter a runtime with low fixed overhead, which is exactly what we claim to be. The bulk test uses 1,000 real sentences instead; passages of several minutes are not yet measured.
  • The bulk comparison is NukeTorch's default bulk mode against MLX generating one request at a time, as it ships. It compares the two as you would run them, not identical scheduling.
  • All three engines were measured on the same machine. Single requests are medians of 60 runs, timed from raw text to finished audio on every engine; the slowest NukeTorch run was 18.80 ms and the fastest MLX run 46.03 ms. No P95 or P99 claim, here or anywhere on this site.
  • Precision: NukeTorch runs FP16 with FP32 accumulation. PyTorch and MLX run FP32, their default configuration, as their packages ship. We did not tune them. NukeTorch's default is qualified on output rather than on bit-for-bit parity. Run at FP32 with MLX’s own phonemes, it is still 2.35× faster per request and 2.58× in bulk.
  • Bulk transcription error rate is 0.13 points higher than MLX’s because of NukeTorch’s own text-to-phoneme step; with Misaki’s phonemes the rates match.

How to use it

macOS on Apple silicon · the same model files and checkpoint revision you use today · licensing undetermined

Benchmark reportReproducing the Kokoro results
MachineApple M4 Max, 40 GPU cores, 128 GB, macOS 26.7
NukeTorch releaseCurrent release build, default configuration
FixtureOne request: “Watch out, driver!”, voice af_heart, 45,600 samples at 24 kHz. Bulk: 1,000 LibriSpeech test-clean sentences
StatisticOne request: 2 rounds of 30 measured runs, median. Bulk: median of 5 passes. No P95 claim.

Exactly what was run

Every engine ran in its default configuration. PyTorch and MLX were not tuned; NukeTorch ran with none of its options set.

PyTorch checkpointhexgrad/Kokoro-82M@f3ff3571791e39611d31c381e3a41a3af07b4987
MLX checkpointprince-canuma/Kokoro-82M@e02c9eada7ce7416798af36b190a8a2dd2ecd566
NukeTorchNative C++/Metal release build, default configuration, its own text-to-phoneme step; bulk mode for the 1,000 sentences
PrecisionPyTorch and MLX in FP32, as their packages ship. NukeTorch FP16 with FP32 accumulation; qualified on output (see the qualification step).
SettingsVoice af_heart, speed 1.0, language a, seed 0, output 24 kHz. PyTorch with disable_complex=True.
EnvironmentPython 3.12.10, PyTorch 2.10.0, MLX 0.32.2, MLX-Audio 0.5.6, Kokoro/Misaki 0.9.4
Bulk workload1,000 LibriSpeech test-clean transcripts in fixed order, sentence case with final punctuation, same list and seed 0 for both engines. Corpus: OpenSLR SLR12, CC BY 4.0.
Benchmark harnessScript revision 325ad9f163ed9e5bec2449f2eec0cfb508b77638. benchmark_framework_voice_workers.py is checked against its recorded SHA-256 before a run counts; the exact executed sources are embedded in each raw receipt.

Steps to reproduce

In order. Qualification comes before anyone should believe the timings: a fast runtime that changes the output has not won anything.

  1. Set up the environment and pin the checkpoints

    Install the package versions above, and download both checkpoints at the exact revisions listed. A different revision is a different benchmark.

  2. Time PyTorch and MLX, full request

    From text received to returned audio. Two rounds per framework, each 2 warmups and then 30 measured repetitions. These are the time to first audio and end-to-end figures.

  3. Time PyTorch and MLX, model only

    Resident neural forward pass after phoneme conversion, tokenization and voice conditioning have been precomputed. Two rounds of 30 measured repetitions per framework. Model load, host PCM copy, output construction and file I/O are excluded.

    # PyTorch and MLX model-only timing
    python kokoro_model_inference_benchmark.py
  4. Time NukeTorch

    Resident public call, from raw text, including NukeTorch’s own text-to-phoneme step, through returned PCM. Two rounds, each 3 warmups and then 30 measured repetitions. Time to first audio and end to end are the wall clock of the public call, the same clock as step 2.

    # NukeTorch resident timing, 3 warmups + 30 measured, two rounds
    nuketorch-audio-time kokoro
  5. Run the bulk job

    1,000 LibriSpeech sentences per pass, five passes per engine, each in a fresh process with the model resident and 2 warmup requests. MLX was measured in one run and NukeTorch in a second. NukeTorch is timed around its batch call, in its default bulk mode, and writes audio files afterwards. MLX generates the requests one after another; its elapsed time is rebuilt from request timestamps and includes the file writes in between. Every pass must return 1,000 non-empty outputs.

  6. Measure memory

    Each engine resident in its own process. The process’s physical footprint, which counts memory used by both the processor and the graphics chip, is sampled from outside the process with the same probe for every engine: peak is the largest sample, and the report also gives the median of the second half of the run. Measured on the single request and over one full bulk pass, in runs separate from the timings, because the probe slows the timed calls.

  7. Qualify the output before trusting the timings

    42 requests in four languages against the reference. Gates, fixed in advance: waveform cosine similarity at least 0.999 (observed at least 0.999998), relative RMSE at most 5% (observed at most 0.20%), identical sample counts, and equal word error rate from automatic transcription on all 36 English requests. Deliberately broken controls must fail the same gates.

  8. Compute the comparisons

    Medians only. Speed: N = reference time ÷ NukeTorch time, shown as N×. Memory: percent less = (1 − NukeTorch ÷ reference) × 100. Recompute real-time factor from raw clocks and actual output duration, not from a package field.

Other configurations

Full precision, like MLXNukeTorch at FP32 with Misaki’s phonemes, one request at a time: 19.88 ms per request (2.35× MLX, 3.10× PyTorch); 60.63 s for 1,000 sentences (2.58× MLX); 2.14% words wrong against 2.12%; 636 MB peak memory against 2,048 MB. NukeTorch’s own FP16 against FP32 with its own phonemes: 18.45 ms against 19.09 ms per request, 2.25% words wrong in both, 998 of 1,000 transcripts identical.

The files behind every number

Each run writes a JSON receipt with the exact invocation, the resolved checkpoint revision, dependency versions, source hashes and the raw per-run clocks the medians were taken from. Raw receipts for the current release are kept in our engineering record; ask and we will share them.

  • Single requests, all engines, 25 Sepfinal-pass-20260925
  • MLX model-only replacement roundfinal-pass-20260925-supplement/kokoro1-mlx-model-r1.json
  • Bulk, 5 passes per enginefinal-pass-20260925/kokoro1000-{nuketorch,mlx}-p{1..5}
  • PyTorch, full request, 17 Sepkokoro-standalone-pytorch-s1-20260917.json
  • MLX, full requestkokoro-standalone-mlx-s1-20260917.json
  • PyTorch, model onlykokoro-model-inference-pytorch-s1-20260918.json
  • MLX, model onlykokoro-model-inference-mlx-s1-20260918.json
  • NukeTorch, 18 Sep releasekokoro-standalone-nuketorch-20260918.json
  • NukeTorch, paired timing and process memorykokoro-a1-k11-pair-20260918.json · kokoro-a1-k11-process-20260918.json
  • NukeTorch, clock cross-checkgpu25-audio-resident.json
  • Output checkscg10-kokoro.json

What this does not show

  • One 1.9-second fixture for single requests. Bulk uses 1,000 real sentences, but nothing of several minutes.
  • The bulk comparison is NukeTorch's default bulk mode against MLX one request at a time. MLX and PyTorch were measured in one run and NukeTorch in a second, not alternated.
  • Medians on one machine. No P95 or P99, no statistical significance claim. No test of several users at once.
  • Transcription error rate is 0.13 points higher than MLX’s on default settings, traced to NukeTorch’s own text-to-phoneme step.
  • A 128 GB M4 Max, where a memory result matters least. Not yet run on a base MacBook Air or an older M1 or M2 machine.
  • No human listening study, besides our own ears. Quality evidence is numerical and automatic transcription.

Text to speech · model 2 of 3

F5-TTS v1 Base

Each clip takes a third less time than with MLX, with identical sampler settings, and a bulk job of 1,000 sentences takes 44% less. It needs about a quarter of the memory.

Mixed precision

Shanghai Jiao Tong University, with Cambridge and Geely · weights CC BY-NC 4.0, non-commercial only · model_1250000, shared reference WAV and transcript, seed 0 · One request: medians of 60 runs · Bulk: 1,000 LibriSpeech sentences, median of 5 passes

MetricNukeTorchMLXPyTorch *vs MLXvs PyTorch *
One request, time to first audio (ms)1,366.532,045.653,229.8333.2% less (1.50×)2.36×
1,000 sentences in bulk (s) §1,955.623,515.26not run44.4% less (1.80×)n/a
Real-time factor, one request (×)1.190.800.531.50×2.28×
Real-time factor, bulk (×)3.251.81not run1.80×n/a
Transcription error rate, bulk2.47%2.52%not runabout the samen/a
Peak memory, one request (MB) †8783,2244,34473% less80% less
Peak memory, 1,000 sentences (MB) †1,741106,189not run98% lessn/a

* PyTorch runs its own default sampler (Euler, 32 steps), not the RK4 settings MLX and NukeTorch share, so its column is not like-for-like. § NukeTorch in its default bulk mode, which schedules the requests itself; MLX generating them one after another, as it ships. Median of five passes; the slowest NukeTorch pass was 1,963.20 s and the fastest MLX pass 3,454.79 s. † Physical footprint sampled from outside the process, the same probe for every engine, in runs separate from the timings. MLX keeps memory it has finished with in its own cache, so its bulk peak reflects its largest working set; its median over the second half of the bulk run was still 105,165 MB, against NukeTorch’s 1,536 MB.

What happens technically

F5 produces audio by running a sampler that evaluates the model 28 times per request. Every one of those evaluations carries the same routing cost, so the overhead is multiplied rather than spread out. Removing it cuts a request from 2,046 ms to 1,367 ms, with the sampler settings held identical on both sides. NukeTorch's timings come from paired runs in balanced order, alternating with a stock MLX worker in the same process; the MLX column is MLX measured on its own, the same day. In bulk the gain grows to 1.80×, because NukeTorch also schedules the requests itself.

What it means for you

F5 is slow on a Mac, so here the time is the point. On the test clip the audio comes back after 1.37 seconds instead of 2.05 under MLX: a third less waiting on every generation. For voice cloning, narration or dubbing done in bulk, 1,000 sentences took 33 minutes instead of 59 under MLX, so work that takes an hour under MLX takes about 33 minutes. Transcription error rates were the same on both: a median of 2.47% and 2.52% across 20,404 words in each of five passes.

On a single request F5 generates audio faster than it plays, at 1.19× real time against MLX's 0.80×, and in bulk at 3.25× real time. And its memory footprint drops from 3.2 GB to 0.9 GB, so it stops needing a well-specified machine. On the 1,000-sentence job the gap is starker: NukeTorch stays at 1.7 GB, while MLX grows to fill 106 GB of this 128 GB Mac.

Identical weights on both sides · F5-TTS 8-bit

Both engines were also run on MLX’s published 8-bit F5-TTS checkpoint, the same file read by both exactly as published, with the same sampler, texts, reference and seeds. NukeTorch is 1.43× faster than MLX on a single request (1,415 ms against 2,022 ms) and 2.01× faster across the 1,000 sentences (30 minutes against 61). Transcription error rates: 2.35% against 2.46%. People pick the 8-bit file to save memory, but on MLX it makes bulk jobs 4% slower than its FP32 model and still needs 2,167 MB for a single request. On NukeTorch the same file runs in 658 MB, and a 1,000-sentence job peaks at 1.6 GB against MLX’s 106 GB.

Does the output still match

Qualified: 36 requests, one to forty words, three reference voices, loud and quiet references. Final mel state differs from stock MLX by a median of 0.13% and at worst 1.44% normalized RMSE, against a 5% gate. Audio transcribes with the same word error rate in every request. The gates were fixed before this build produced its first output.

Nuance
  • NukeTorch and MLX use the same sampler settings.
  • The bulk comparison is NukeTorch's default bulk mode against MLX one request at a time, as it ships. The median of five passes is shown.
  • Timing boundaries differ slightly: MLX and PyTorch request timing includes reading the reference and preparing the text; NukeTorch starts at its speak call.
  • Precision: NukeTorch runs FP16 with FP32 accumulation; MLX and PyTorch run FP32 reference weights. With the same 8-bit checkpoint on both sides, NukeTorch is still 1.43× faster per request and 2.01× in bulk.
  • Memory is the peak physical footprint, NukeTorch in FP16 against MLX and PyTorch in FP32, measured with the same outside probe. MLX’s 106 GB bulk peak partly reflects memory it keeps cached after use; on a smaller Mac it would likely use less and run slower.
  • The whole clip returns at once, so the listener waits the full 1.4 seconds before hearing anything. We have not measured sentence-by-sentence streaming and make no claim about live conversation.
  • Output lengths differ slightly between samplers: 39,181 samples for MLX and NukeTorch, 40,704 for PyTorch. Real-time factor is computed per request from actual output duration.
  • Text outside English, or a reference that is not mono PCM16, falls back to an MLX worker rather than failing.

How to use it

macOS on Apple silicon · the same model files and checkpoint revision you use today · licensing undetermined

Benchmark reportReproducing the F5-TTS results
MachineApple M4 Max, 40 GPU cores, 128 GB, macOS 26.7
NukeTorch releaseaudio20260925, measured 25 Sep 2026, default configuration
FixtureOne request: “Watch out, driver!”, packaged reference WAV, seed 0. Bulk: 1,000 LibriSpeech test-clean sentences
RunsNukeTorch: ABBA then BAAB paired series, 3 warmups + 30 measured per mode. MLX and PyTorch: 2 rounds of 30. Bulk: median of 5 passes

Start with the single-request medians: 1,366.53 ms for NukeTorch against 2,045.65 ms for MLX, 60 runs each, same day, same sampler.

Exactly what was run

PyTorch checkpointSWivid/F5-TTS@84e5a410d9cead4de2f847e7c9369a6440bdfaca, variant F5TTS_v1_Base, model_1250000
MLX workerThe same file NukeTorch reads: lucasnewman/f5-tts-mlx snapshot 2d719cadec8fd3887c8599475a5f65924d523654
NukeTorchInstalled release audio20260925. The whole request runs natively in the calling process, with no Python or MLX. Text outside English, or a reference that is not mono PCM16 WAV, uses the MLX worker instead.
Sampler, MLX and NukeTorchRK4, 8 time points (28 flow evaluations), CFG 2, sway −1. Core tensor [1,650,100].
Sampler, PyTorchIts production default: Euler, 32 intervals, CFG 2, sway −1. Core tensor [1,625,100]. Not identical numerical work, so the PyTorch column should be discounted.
PrecisionPyTorch and MLX in FP32, as their packages run on this machine. NukeTorch FP16 with FP32 accumulation; qualified on output (step 2).
EnvironmentPython 3.12.10, f5-tts-mlx 0.2.6 with MLX 0.31.2, PyTorch 2.10.0
Bulk workload1,000 LibriSpeech test-clean transcripts in fixed order, same list and seed 0 for both engines, packaged English reference. Corpus: OpenSLR SLR12, CC BY 4.0.

Steps to reproduce

From the NukeTorch repository root. The harness first proves that its library and kernels are byte-identical to the installed build, so you are timing what ships.

  1. Build and verify

    Fresh Release build. The recording script checks that the harness library and Metal kernels match the installed build byte for byte before any run counts.

  2. Qualify the output

    36 requests against stock MLX: texts of 1 to 40 words, three reference voices, loud and quiet references. Gates fixed before this build's first output existed: final mel state within 5% normalized RMSE (observed median 0.13%, worst 1.44%), and equal word error rate from automatic transcription in every request. Known-fail controls must fail.

    python3 bench/f5_stage_gate.py qualify --library … --metallib … --output …
  3. Record the same-run control

    One benchmark process alternates a resident stock-MLX worker and the native path: ABBA, then the designated BAAB. 3 warmups and 30 measurements per mode, with a 37-check audit of every measured request. It then records first call and process memory for the native path alone.

    # RUN_ID must be unused
    bench/f5_a2_record.sh RUN_ID
  4. Apply the drift rule

    A BAAB run counts only if its same-run stock MLX model inference lands within 2% of the saved 2,016.45 ms. Otherwise the machine idles for 7 minutes and the run repeats. This keeps a hot or busy machine from flattering either side. In the current series the control stayed within the limit (+1.7% and −1.3%), so no run had to be repeated.

  5. Run the bulk job

    1,000 LibriSpeech sentences per pass, five passes per engine, each in a fresh process with the model resident and 2 warmup requests. NukeTorch runs in its default bulk mode and is timed around its batch call; MLX generates the requests one after another, with elapsed time rebuilt from request timestamps. Score the outputs with Whisper large-v3-turbo against the corpus text.

  6. Write the report
    python3 bench/f5_native_report_update.py --audit … --rss … --in-process … --public-summary …
  7. Optional: rerun the original model-only baseline

    The earlier three-framework baseline, with an uninstrumented audio control per framework. 2 solver warmups and 5 measured solver replays per framework. f5_a2_delivery_run.py reruns time to first audio, end to end and real-time factor on the same fixture.

    python3 bench/f5_a2_run.py --run-id UNUSED_RUN_ID --native build/FRESH_BUILD_ID/nuketorch-f5-a2-native
    python3 bench/f5_a2_audit.py --run-id UNUSED_RUN_ID

How each number is measured

Time to first audioStarts before the request API call, stops when the full audio returns. The API returns one buffer, so time to first audio equals end to end. That is the delivery mode, not a streaming result.
Real-time factorMedian of per-request output duration ÷ end to end. Output: 39,181 samples for MLX and NukeTorch, 40,704 for PyTorch, at 24 kHz.
Peak memoryThe process’s physical footprint on the single-request fixture, model resident, sampled from outside the process with the same probe for NukeTorch and MLX. Measured in separate runs from the timings.
Bulk job1,000 requests divided by the batch interval gives throughput; real-time factor is generated audio divided by the interval. Transcription error rate compares Whisper large-v3-turbo transcripts of the output with the corpus text, 20,404 words.

Other configurations

8-bit, same checkpointMLX’s published model_v1_8b, read by both engines as published (same file, same SHA-256): 1,415.17 ms against 2,021.68 ms per request, median of 60 calls (1.43×); 1,815.52 s against 3,650.81 s for 1,000 sentences, median of 5 passes (2.01×); 2.35% against 2.46% words wrong; peak memory 658 MB against 2,167 MB on a single request and 1,563 MB against 106,189 MB over the bulk job. Same-process MLX control within 2% (+0.9% and −1.6%).

The files behind every number

Raw receipts for the current release are kept in our engineering record; ask and we will share them. The original baseline receipts are listed below.

  • Single requests, all engines, 25 Sepfinal-pass-20260925
  • NukeTorch paired series, 25 Sepfinal-f5-a2-final-pass-20260925-{abba,baab1}-audit.json
  • Bulk, pass 1, and transcription scoresfinal-pass-20260925/f5-1000-{nuketorch,mlx}-p1 · summary.json
  • PyTorch, model only, 18 Sepf5-a2-production-20260918-r4-pytorch.json
  • MLX, model onlyf5-a2-production-20260918-r4-mlx.json
  • NukeTorch worker, model onlyf5-a2-production-20260918-r4-nuketorch.json
  • Model-only commands and controlsf5-a2-production-20260918-r4-manifest.json
  • Model-only 31-check auditf5-a2-production-20260918-r4-audit.md
  • Delivery timing, three frameworksf5-a2-delivery-20260918-r1-{pytorch,mlx,nuketorch}.json
  • Delivery commands and 21-check auditf5-a2-delivery-20260918-r1-manifest.json · f5-a2-delivery-20260918-r1-audit.md
  • Standalone timingaudio-standalone-f5-pytorch-20260917.json · audio-standalone-f5-mlx-20260918.json

What this does not show

  • One timing fixture for single requests; bulk uses 1,000 sentences, five passes per engine. No statistical significance or general speed claim.
  • No streaming test. The whole clip returns at once.
  • The PyTorch column compares different samplers and should be discounted; the MLX comparison does not have this problem.
  • No P95 or P99. No test of several users at once.
  • Bulk memory was measured over one pass per engine; PyTorch was not run in bulk.
  • No NVIDIA comparison, including against TensorRT-LLM.
  • No human listening study, besides our own ears. Quality evidence is numerical and automatic transcription.

Text to speech · model 3 of 3

Fish Speech S2 Pro

The heaviest model, and the cleanest comparison, with both engines at the same precision. 1.2× faster than MLX on a single request, which is what the dispatch argument predicts, and 4.4× on a bulk job.

Same precision

Fish Audio · Fish Audio Research License · Built with Fish Audio · both engines on the checkpoint's BF16 weights, no quantization · One request: equal work at 33 frames, medians of 60 runs · Bulk: 1,000 LibriSpeech sentences, median of 5 passes

MetricNukeTorchMLXvs MLX
One request, full clip, 33 frames (ms)1,248.351,509.851.21×
Time per generated frame (ms)37.8345.751.21×
Real-time factor, one request (×)1.231.021.21×
1,000 sentences in bulk (s) §1,445.996,373.304.41×
Real-time factor, bulk (×)4.651.064.41×
Transcription error rate, bulk2.00%1.94%about the same
Peak memory, one request (MB)12,45815,20618% less
Peak memory, 1,000 sentences (MB)30,720108,47972% less

§ NukeTorch in its default bulk mode, which schedules the requests itself; MLX generating them one after another, as it ships. Median of five passes; the slowest NukeTorch pass was 1,544.39 s and the fastest MLX pass 6,346.40 s.

What happens technically

S2 Pro is the heaviest model measured here, and on a single request it shows the smallest gain: 1.21×, with both engines reading the same BF16 weights and doing the same 33 frames of work. The larger the model, the more of the clock is genuine arithmetic and the less of it is routing. If the single-request gain had grown with model size, that would be evidence against us.

The bulk result is a different effect. NukeTorch's default bulk mode schedules the 1,000 requests itself, while MLX, as shipped, generates them one after another. On a model this heavy, that scheduling is worth 4.4×.

What it means for you

On a single request, 1.21× faster means about one hour of compute saved in every six. In bulk, 1,000 sentences took 24 minutes instead of 1 hour 46 minutes under MLX, with transcription error rates of 2.00% and 1.94%, a difference too small to matter. For a business whose margin is the gap between what a minute of audio costs to produce and what it sells for, that matters, and it does not require changing the model, the weights or the output.

Nuance
  • MLX is the only comparison here. There is no usable PyTorch baseline for this model on a Mac.
  • Same precision on both sides: NukeTorch reads the checkpoint's BF16 weights exactly and accumulates in FP32; MLX runs the BF16 model as shipped. Neither uses 8-bit weights in the current tables.
  • Fish samples its audio, so the two engines produce different valid renderings. For the single-request test, NukeTorch was capped at the 33 frames MLX produced, so both do the same work.
  • The bulk comparison is NukeTorch's default bulk mode against MLX one request at a time, as it ships. NukeTorch's first pass took 6.7 to 8.9% longer than the other four, for a cause not established; a sixth pass on an idle machine (1,424 s) matched passes 2 to 5 and is not counted. In that first pass, two of the 1,000 requests reached the length limit rather than ending on their own; in the other four, every request ended on its own.

How to use it

macOS on Apple silicon · licensing undetermined

Anticipating

Questions about these results

Including the ones that do not flatter us. Questions about licensing and NVIDIA are on the home page.

What have you not measured?

This list is longer than the results, which is the correct ratio for a runtime this young.

  • NVIDIANothing. Everything here is Apple silicon, and we will not ask for a place in a runtime comparison table until that changes.
  • StreamingEvery timing returns the whole clip at once. We have not measured sentence-by-sentence streaming, so we make no claim about live conversation.
  • Long passagesSingle-request results use one short sentence. The bulk tests use 1,000 real sentences, but nothing of several minutes. Long text may shrink the gain, and we will not claim otherwise until we have run it.
  • Other MacsEverything here is a 128 GB M4 Max, the machine where a memory result matters least. A base MacBook Air and an older M1 or M2 machine are next.
  • Tail latencySingle requests are medians of 60 runs, with the fastest and slowest run reported. No P95, no P99.
  • Dispatch, isolatedWe have not published a measurement that separates routing time from arithmetic. Every figure is end to end, which is why the schematic is drawn rather than plotted.
  • ConcurrencyThe bulk tests measure one large job. How many simultaneous users one machine can serve is not yet measured.
  • Listening testsNo human listening study, besides our own ears. The quality evidence is numerical and automatic transcription only.
  • Sustained loadNo thermal data. Ten minutes of continuous generation on a fanless machine may look nothing like our 128 GB M4 Max with fans.
  • Cold startModel load time is excluded. For a desktop app it is often the number a user actually feels.
  • EnergyWatt-hours per minute of audio. Nobody in this space publishes it, and we should be the ones to start.
Does the output change?

Same model files, same checkpoint revision, same sampler defaults. Each model section carries its own evidence: waveform similarity and word error rate for Kokoro, mel state and word error rate for F5, with the gates fixed before the builds produced their first output and deliberately broken controls proving the gates bite.

What we do not claim is that it sounds identical. The checks are automated, and beyond our own ears nobody has run a listening study yet. NukeTorch runs at half precision (FP16, or the checkpoint's own BF16 for Fish) and accumulates in full precision; beyond that we have not published its internals.

How do I check any of this myself?

Every model has a benchmark report with the pinned checkpoints, the environment, the procedure, and the limits of each number. For now, we run private demos on request.

Parakeet transcribes about 3× faster. Whisper up to 2.5× faster. With the same accuracy.

Compared with MLX, the fastest alternative on a Mac. Same model files, no changes to your code, the same word error rate on 1,000 real recordings, and a fraction of the memory.

Parakeet TDT 0.6B v3 · one clip and in bulk
One clip
20 ms
vs 47 ms · 2.3×
1,000 clips
14 s
vs 48 s · 3.4×

Against MLX. A 6.4-second utterance, and two hours of real recordings.

Whisper · 1,000 clips
Tiny
12.6 s
vs 32.0 s · 2.5×
large-v3-turbo
135 s
vs 239 s · 1.8×

Against MLX, with FP16 weights on both sides.

Peak memory · 1,000 clips
Parakeet
2.3 GB
vs 45.8 GB
Whisper turbo
2.5 GB
vs 5.1 GB

The process’s physical footprint, NukeTorch against MLX, measured the same way for both.

Measured on one Apple M4 Max with the same model files the authors publish. One clip: 6.4 seconds of English speech, medians of 60 calls. Bulk: 1,000 LibriSpeech recordings, medians of 5 passes, with NukeTorch in its default bulk mode and MLX one clip at a time. No comparison with WhisperKit or NVIDIA yet. What we have not measured. Why it is faster: how NukeTorch works.

Why it matters

The difference

The same results land differently depending on what you build.

48 s → 14 s

You transcribe in bulk

Meeting notes, call recordings, podcast archives. Two hours of recorded speech transcribed by Parakeet in 14 seconds instead of 48 under MLX, with the same word error rate, in 2.3 GB of memory instead of 45.8 GB.

Parakeet, 1,000 real recordings
47 ms → 20 ms

You run a local voice assistant

Listening is the first wait in every turn. Parakeet turns a 6.4-second utterance into text in 20 ms instead of 47, before the language model and the voice even start. Pair it with the text to speech results for both ends of the conversation.

Parakeet, one clip
239 s → 135 s

You ship Whisper in a Mac app

The same Whisper weights at the same precision, no retraining, 1.8× faster on large-v3-turbo across 1,000 recordings and in half the memory. Whisper Tiny is 2.5× faster.

Whisper large-v3-turbo, 1,000 real recordings

Evidence, in one table

Three models, one Mac

MLX is the comparison we hold ourselves to, because it is the fastest alternative on a Mac. PyTorch appears in each model's table because it is the reference the model authors publish against.

Model One clip vs MLX 1,000 clips vs MLX Memory in bulk vs MLX Words wrong, NukeTorch vs MLX Evidence
Parakeet TDT 0.6B v3 2.3× faster 3.4× faster 95% less 1.77% vs 1.75% Mixed precision
Whisper large-v3-turbo 1.2× faster 1.8× faster 52% less 3.18% vs 3.20% Same precision
Whisper Tiny 2.0× faster 2.5× faster 70% less 7.04% vs 7.04% Same precision
One clip

A 6.4-second English utterance, timed from audio in memory to the final transcript. Median of 60 calls.

1,000 clips

Real recordings from the LibriSpeech corpus, 7,389 seconds of speech. NukeTorch runs in its default bulk mode, which schedules the clips itself; MLX transcribes them one after another, as it ships.

Word error rate

The share of words transcribed wrong against the human transcript. Lower is better; equal numbers mean the speed came without an accuracy cost.

Speech to text · model 1 of 3

Parakeet TDT 0.6B v3

About 3× faster than MLX: 2.3× on a single clip and 3.4× across 1,000 recordings, with the same word error rate. In bulk it needs 95% less memory.

Mixed precision

NVIDIA · CC BY 4.0 · 25 languages · One clip: 6.4 seconds of English speech, median of 60 calls · Bulk: 1,000 LibriSpeech recordings, 7,389 seconds of speech, median of 5 passes

MetricNukeTorchMLXPyTorchvs MLXvs PyTorch
One clip, request time (ms)20.0946.9085.792.33×4.27×
1,000 clips (s) §14.1448.05not run3.40×n/a
Clips per second, bulk70.7220.81not run3.40×n/a
Real-time factor, one clip (×)318.6136.574.62.33×4.27×
Real-time factor, bulk (×)522.6153.8not run3.40×n/a
Word error rate, 1,000 clips1.77%1.75%not runabout the samen/a
Peak memory, one clip (MB)1,7413,1743,89145% less55% less
Peak memory, 1,000 clips (MB) §2,25345,773not run95% lessn/a

§ 1,000 LibriSpeech test-clean recordings, 7,389 seconds of speech, median of five passes per engine. NukeTorch runs in its default bulk mode, which schedules the clips itself; MLX transcribes them one after another, as it ships. PyTorch was not run in bulk. Memory is the process’s physical footprint sampled from outside, the same probe for every engine, in runs separate from the timings. MLX keeps memory it has finished with in its own cache, so its bulk peak reflects its largest working set.

What happens technically

Parakeet's decoder works through the audio step by step, emitting up to ten symbols per slice of sound, and each step is a handful of small operations that each have to be routed before they run. That is the shape of work NukeTorch is built for: 2.3× faster than MLX on a single clip, and 3.4× across 1,000 recordings, where NukeTorch also schedules the clips itself. Memory is the other half of the result. Over the bulk job MLX’s footprint climbs to 45.8 GB, partly because it holds on to memory it has finished with; NukeTorch stays at 2.3 GB.

What it means for you

Two hours of recorded speech, 1,000 clips, transcribed in 14 seconds instead of 48 under MLX, with the same word error rate, in 2.3 GB of memory instead of 45.8 GB. For meeting notes, call recordings or podcast archives, that is a batch job that fits on an ordinary Mac.

In a local voice assistant, listening is the first wait in every turn. A 6.4-second utterance becomes text in 20 ms instead of 47 ms, before the language model and the voice even start.

Nuance
  • Precision: NukeTorch runs FP16. MLX and PyTorch use the FP32 weights as shipped, with MLX computing the audio features in BF16. The word error rates match, but it is not a same-precision comparison.
  • The bulk comparison is NukeTorch’s default bulk mode against MLX one clip at a time, as it ships.
  • MLX keeps memory it has finished with in its own cache, so its 45.8 GB bulk peak reflects its largest working set rather than a steady need. Its median over the second half of the run was still 41.5 GB, against NukeTorch’s 2.3 GB.
  • The one-clip sample was generated with Kokoro, not recorded. The 1,000-clip test uses real recordings.
  • English only, although Parakeet v3 supports 25 languages.

How to use it

macOS on Apple silicon · the same model files and checkpoint revision you use today · licensing undetermined

Benchmark reportReproducing the Parakeet results
MachineApple M4 Max, 40 GPU cores, 128 GB, macOS 26.7
Reportv1.0
One clipMedian of 60 timed calls per engine (2 rounds of 30)
1,000 clipsMedian of 5 passes per engine; PyTorch not run in bulk

Exactly what was run

ModelNVIDIA Parakeet TDT 0.6B v3, 600M parameters, 25 languages
CheckpointsNVIDIA nvidia/parakeet-tdt-0.6b-v3@541d1f99c6b0c3cd0b11a95167540bb8edefd82b; MLX Community mlx-community/parakeet-tdt-0.6b-v3@ed2b7e8c15f9aaa0b5772e2efb986255eaef7e15
PrecisionNukeTorch: FP16 weights and activations, FP32 accumulation. MLX and PyTorch: FP32 weights as shipped, MLX with BF16 audio features.
DecodingGreedy TDT decoding, English, up to 10 symbols per frame, on all three engines
Versionsmlx-audio 0.5.6 on MLX 0.32.2; Transformers 5.17.0 on PyTorch 2.10.0, MPS device
MachineApple M4 Max, 40-core GPU, 128 GB unified memory, macOS 26.7, on AC power, no other workload during timing
NukeTorchNative C++ and Metal, release build; converts the original checkpoint to its own FP16 weight format; bulk mode used for the 1,000 clips
One-clip sample5.6 seconds of English speech generated with Kokoro (voice af_heart), resampled to 16 kHz mono and zero-padded to 6.4 seconds, already in memory
1,000-clip sampleLibriSpeech test-clean, the first 1,000 recordings in ID order after skipping clips over 30 seconds; 7,389 seconds in total. FLAC converted losslessly to 16-bit WAV; reference transcripts from the corpus. OpenSLR SLR12, CC BY 4.0.

How each number is measured

One clipTimer from audio samples already in memory, through features, encoder and greedy decoding, to the text out. Model loaded and resident; model loading and file reading are excluded. Two rounds of 30 timed calls per engine, pooled: NukeTorch after 3 warmups per round, MLX and PyTorch after 2.
1,000 clipsFive passes per engine, each in its own process, two warmup clips before each. NukeTorch receives the list through its bulk mode, which schedules the clips itself; MLX transcribes one clip after another in a Python loop. Wall time is the computer's clock around the whole pass. Throughput is clips divided by wall time; real-time factor is 7,389 seconds of audio divided by wall time.
MemoryThe process's physical footprint, sampled from outside the process with the same probe for every engine. Peak is the largest sample; the report also gives the median of the second half of the run. Measured in separate runs, because the probe slows the timed calls.
Word error rateSubstituted, missing and added words divided by the 20,404 words in the human transcript, after lowercasing and removing punctuation, the same scorer for both engines. The one-clip result is only a check; the 1,000-clip figure is the real result.

What this does not show

  • Different precision: NukeTorch FP16 against FP32 weights.
  • One machine. No P95 or P99, no statistical significance claim.
  • Clips of up to 30 seconds; no hour-long files. English read speech only.
  • No streaming: these runtimes do not expose partial transcripts, so there is no time-to-first-word figure.
  • The bulk comparison is NukeTorch’s own scheduling against MLX one clip at a time, not matched concurrency.
  • No comparison yet with other Mac runtimes such as WhisperKit or whisper.cpp, or with NVIDIA.

Speech to text · model 2 of 3

Whisper large-v3-turbo

The heaviest speech-to-text model here and the cleanest comparison, with FP16 weights on every engine: 1.2× faster than MLX on a single clip, 1.8× across 1,000 recordings, in half the memory.

Same precision

OpenAI · MIT · multilingual · One clip: 6.4 seconds of English speech, median of 60 calls · Bulk: 1,000 LibriSpeech recordings, median of 5 passes

MetricNukeTorchMLXPyTorchvs MLXvs PyTorch
One clip, request time (ms)193.21228.57368.991.18×1.91×
1,000 clips (s) §134.81239.18not run1.77×n/a
Clips per second, bulk7.424.18not run1.77×n/a
Real-time factor, one clip (×)33.128.017.31.18×1.91×
Real-time factor, bulk (×)54.830.9not run1.77×n/a
Word error rate, 1,000 clips3.18%3.20%not runabout the samen/a
Peak memory, one clip (MB)1,8432,7653,27733% less44% less
Peak memory, 1,000 clips (MB) §2,4585,120not run52% lessn/a

§ 1,000 LibriSpeech test-clean recordings, 7,389 seconds of speech, median of five passes per engine. NukeTorch runs in its default bulk mode, which schedules the clips itself; MLX transcribes them one after another, as it ships. PyTorch was not run in bulk. Memory is the process’s physical footprint sampled from outside, the same probe for every engine, in runs separate from the timings. MLX keeps memory it has finished with in its own cache, so its bulk peak reflects its largest working set.

What happens technically

large-v3-turbo is the heaviest speech-to-text model on this page, and on a single clip it shows the smallest gain, 1.18×. The larger the model, the more of the clock is genuine arithmetic and the less of it is routing, so this is what the argument predicts; if the gain had grown with model size, that would count against us. All three engines run FP16 weights, which makes this the most like-for-like result on the page. Across 1,000 recordings, with NukeTorch scheduling the clips itself, the gain is 1.77×.

What it means for you

If you ship Whisper in a Mac app, this is the result that transfers directly: the same weights, the same precision, no retraining, and 1,000 recordings transcribed in 2¼ minutes instead of 4, in 2.5 GB of memory instead of 5.1 GB. Accuracy is unchanged: 3.18% of words wrong against MLX’s 3.20%.

Nuance
  • FP16 weights on every engine.
  • Decoding settings differ slightly: all three decode greedily in English, but MLX also produces its default timestamp tokens, PyTorch caps output at 64 new tokens, and NukeTorch produces no timestamps.
  • The bulk comparison is NukeTorch’s default bulk mode against MLX one clip at a time, as it ships.
  • The one-clip sample was generated with Kokoro, not recorded. The 1,000-clip test uses real recordings.
  • The comparison buyers will ask for, WhisperKit on Mac and faster-whisper on NVIDIA, has not been run yet.

How to use it

macOS on Apple silicon · the same model files and checkpoint revision you use today · licensing undetermined

Benchmark reportReproducing the Whisper large-v3-turbo results
MachineApple M4 Max, 40 GPU cores, 128 GB, macOS 26.7
Reportv1.0
One clipMedian of 60 timed calls per engine (2 rounds of 30)
1,000 clipsMedian of 5 passes per engine; PyTorch not run in bulk

Exactly what was run

ModelOpenAI Whisper large-v3-turbo, 809M parameters, multilingual
CheckpointsOpenAI openai/whisper-large-v3-turbo@41f01f3fe87f28c78e2fbf8b568835947dd65ed9; MLX Community mlx-community/whisper-large-v3-turbo@a4aaeec0636e6fef84abdcbe3544cb2bf7e9f6fb
PrecisionFP16 weights on all three engines; NukeTorch accumulates in FP32
DecodingEnglish, greedy. NukeTorch and PyTorch without timestamps; MLX as shipped, with its default timestamp tokens decoded but not counted. MLX: temperature 0, previous-text conditioning off. PyTorch: eager attention, at most 64 new tokens.
Versionsmlx-whisper 0.4.3 on MLX 0.32.2; Transformers 5.3.0 on PyTorch 2.10.0, MPS device
MachineApple M4 Max, 40-core GPU, 128 GB unified memory, macOS 26.7, on AC power, no other workload during timing
NukeTorchNative C++ and Metal, release build; converts the original checkpoint to its own FP16 weight format; bulk mode used for the 1,000 clips
One-clip sample5.6 seconds of English speech generated with Kokoro (voice af_heart), resampled to 16 kHz mono and zero-padded to 6.4 seconds, already in memory
1,000-clip sampleLibriSpeech test-clean, the first 1,000 recordings in ID order after skipping clips over 30 seconds; 7,389 seconds in total. FLAC converted losslessly to 16-bit WAV; reference transcripts from the corpus. OpenSLR SLR12, CC BY 4.0.

How each number is measured

One clipTimer from audio samples already in memory, through features, encoder and greedy decoding, to the text out. Model loaded and resident; model loading and file reading are excluded. Two rounds of 30 timed calls per engine, pooled: NukeTorch after 3 warmups per round, MLX and PyTorch after 2.
1,000 clipsFive passes per engine, each in its own process, two warmup clips before each. NukeTorch receives the list through its bulk mode, which schedules the clips itself; MLX transcribes one clip after another in a Python loop. Wall time is the computer's clock around the whole pass. Throughput is clips divided by wall time; real-time factor is 7,389 seconds of audio divided by wall time.
MemoryThe process's physical footprint, sampled from outside the process with the same probe for every engine. Peak is the largest sample; the report also gives the median of the second half of the run. Measured in separate runs, because the probe slows the timed calls.
Word error rateSubstituted, missing and added words divided by the 20,404 words in the human transcript, after lowercasing and removing punctuation, the same scorer for both engines. The one-clip result is only a check; the 1,000-clip figure is the real result.

What this does not show

  • One machine. No P95 or P99, no statistical significance claim.
  • Clips of up to 30 seconds; no hour-long files. English read speech only.
  • No streaming: these runtimes do not expose partial transcripts, so there is no time-to-first-word figure.
  • The bulk comparison is NukeTorch’s own scheduling against MLX one clip at a time, not matched concurrency.
  • No comparison yet with other Mac runtimes such as WhisperKit or whisper.cpp, or with NVIDIA.

Speech to text · model 3 of 3

Whisper Tiny

The smallest model, where routing is most of the clock: 2.0× faster than MLX on a single clip, 2.5× across 1,000 recordings, and 11× faster than PyTorch.

Same precision

OpenAI · Apache 2.0 · multilingual · One clip: 6.4 seconds of English speech, median of 60 calls · Bulk: 1,000 LibriSpeech recordings, median of 5 passes

MetricNukeTorchMLXPyTorchvs MLXvs PyTorch
One clip, request time (ms)13.1926.62151.582.02×11.49×
1,000 clips (s) §12.6332.03not run2.54×n/a
Clips per second, bulk79.1631.22not run2.54×n/a
Real-time factor, one clip (×)485.2240.442.22.02×11.49×
Real-time factor, bulk (×)584.9230.7not run2.54×n/a
Word error rate, 1,000 clips7.04%7.04%not runthe samen/a
Peak memory, one clip (MB)3566941,63849% less78% less
Peak memory, 1,000 clips (MB) §9153,072not run70% lessn/a

§ 1,000 LibriSpeech test-clean recordings, 7,389 seconds of speech, median of five passes per engine. NukeTorch runs in its default bulk mode, which schedules the clips itself; MLX transcribes them one after another, as it ships. PyTorch was not run in bulk. Memory is the process’s physical footprint sampled from outside, the same probe for every engine, in runs separate from the timings. MLX keeps memory it has finished with in its own cache, so its bulk peak reflects its largest working set.

What happens technically

Tiny is small enough that very little of a request is arithmetic, so removing the routing matters most here. Against PyTorch, which routes every operation through its general-purpose dispatcher, the gap is 11.5×. Against MLX, which already closes most of that gap, it is 2.0× on a single clip and 2.5× across 1,000 recordings, with FP16 weights on every engine.

What it means for you

Tiny is the model people pick when speed and footprint matter more than accuracy: captions, voice commands, on-device search. Here 1,000 recordings take under 13 seconds instead of 32 under MLX, in 915 MB of memory instead of 3.1 GB, at exactly the same word error rate.

Nuance
  • FP16 weights on every engine.
  • The bulk comparison is NukeTorch’s default bulk mode against MLX one clip at a time, as it ships.
  • The one-clip sample was generated with Kokoro, not recorded. The 1,000-clip test uses real recordings.
  • Tiny's accuracy is modest on any engine: about 7% of words wrong on this corpus.

How to use it

macOS on Apple silicon · the same model files and checkpoint revision you use today · licensing undetermined

Benchmark reportReproducing the Whisper Tiny results
MachineApple M4 Max, 40 GPU cores, 128 GB, macOS 26.7
Reportv1.0
One clipMedian of 60 timed calls per engine (2 rounds of 30)
1,000 clipsMedian of 5 passes per engine; PyTorch not run in bulk

Exactly what was run

ModelOpenAI Whisper Tiny, 39M parameters, multilingual
CheckpointsOpenAI openai/whisper-tiny@169d4a4341b33bc18d8881c4b69c2e104e1cc0af; MLX Community mlx-community/whisper-tiny@78c52ab98ca87f570bc57ad852e15ef7060f9f76
PrecisionFP16 weights on all three engines; NukeTorch accumulates in FP32
DecodingEnglish, greedy, no timestamps, on all three engines. MLX: temperature 0, previous-text conditioning off. PyTorch: eager attention, at most 64 new tokens.
Versionsmlx-whisper 0.4.3 on MLX 0.32.2; Transformers 5.3.0 on PyTorch 2.10.0, MPS device
MachineApple M4 Max, 40-core GPU, 128 GB unified memory, macOS 26.7, on AC power, no other workload during timing
NukeTorchNative C++ and Metal, release build; converts the original checkpoint to its own FP16 weight format; bulk mode used for the 1,000 clips
One-clip sample5.6 seconds of English speech generated with Kokoro (voice af_heart), resampled to 16 kHz mono and zero-padded to 6.4 seconds, already in memory
1,000-clip sampleLibriSpeech test-clean, the first 1,000 recordings in ID order after skipping clips over 30 seconds; 7,389 seconds in total. FLAC converted losslessly to 16-bit WAV; reference transcripts from the corpus. OpenSLR SLR12, CC BY 4.0.

How each number is measured

One clipTimer from audio samples already in memory, through features, encoder and greedy decoding, to the text out. Model loaded and resident; model loading and file reading are excluded. Two rounds of 30 timed calls per engine, pooled: NukeTorch after 3 warmups per round, MLX and PyTorch after 2.
1,000 clipsFive passes per engine, each in its own process, two warmup clips before each. NukeTorch receives the list through its bulk mode, which schedules the clips itself; MLX transcribes one clip after another in a Python loop. Wall time is the computer's clock around the whole pass. Throughput is clips divided by wall time; real-time factor is 7,389 seconds of audio divided by wall time.
MemoryThe process's physical footprint, sampled from outside the process with the same probe for every engine. Peak is the largest sample; the report also gives the median of the second half of the run. Measured in separate runs, because the probe slows the timed calls.
Word error rateSubstituted, missing and added words divided by the 20,404 words in the human transcript, after lowercasing and removing punctuation, the same scorer for both engines. The one-clip result is only a check; the 1,000-clip figure is the real result.

What this does not show

  • One machine. No P95 or P99, no statistical significance claim.
  • Clips of up to 30 seconds; no hour-long files. English read speech only.
  • No streaming: these runtimes do not expose partial transcripts, so there is no time-to-first-word figure.
  • The bulk comparison is NukeTorch’s own scheduling against MLX one clip at a time, not matched concurrency.
  • No comparison yet with other Mac runtimes such as WhisperKit or whisper.cpp, or with NVIDIA.

Anticipating

Questions about these results

Including the ones that do not flatter us. Questions about licensing and NVIDIA are on the home page.

What have you not measured?

This list is longer than the results, which is the correct ratio for a runtime this young.

  • Other Mac runtimesNo comparison yet with WhisperKit, the usual Whisper runtime in Mac and iPhone apps, or with whisper.cpp.
  • NVIDIANothing. The comparison speech-to-text buyers expect there is faster-whisper, and we will not claim anything about servers until we have run it.
  • Matched bulk runsThe bulk results compare NukeTorch’s own scheduling with MLX one clip at a time. We have not yet published a one-clip-at-a-time NukeTorch run for these models.
  • StreamingThese runtimes do not expose partial transcripts, so there is no time-to-first-word figure. Every result is the full transcript.
  • Long audioClips of up to 30 seconds. No hour-long recordings.
  • LanguagesEnglish only, although Parakeet and Whisper are multilingual.
  • Precision on ParakeetNukeTorch runs FP16 against the FP32 weights MLX and PyTorch ship with. Both Whisper models use FP16 weights everywhere.
  • Tail latencyMedians of 60 calls, with the fastest and slowest reported. No P95, no P99.
  • Real one-clip audioThe one-clip sample was generated with Kokoro. The 1,000-clip test uses real recordings.
  • Other MacsEverything here is a 128 GB M4 Max. A base MacBook Air and an older M1 or M2 machine are next.
  • Cold startModel load time is excluded. For a desktop app it is often the number a user actually feels.
Does the transcript change?

Same model files and checkpoint revisions. On the 1,000 recordings, NukeTorch’s word error rate is within 0.02 percentage points of MLX on all three models: 1.77% against 1.75% for Parakeet, 3.18% against 3.20% for Whisper large-v3-turbo, and 7.04% against 7.04% for Whisper Tiny. On the one-clip test, all three engines returned the same transcript on every call.

How do I check any of this myself?

Every model has a benchmark report with the pinned checkpoints, the environment, the procedure, and the limits of each number. For now, we run private demos on request.

Get in touch

A demo, a pilot, or just a question — write to us.