WakeHuBERT tiny

A 0.64M-parameter streaming speech feature extractor for wake-word detection, distilled from HuBERT-base. It turns 16 kHz audio into 128-dimensional features at 50 frames per second, so a small classifier trained on those features, from synthetic speech alone, can detect a wake word typed as text. It is the extractor behind wakeforge's wakehubert featurizer.

Using it

The ONNX graph takes waveform [batch, samples] (float32, 16 kHz, -1..1) and returns features [batch, samples // 320, 128]. The extractor is strictly causal: frame t depends only on audio before sample 320·(t+1), within a receptive field of 2.5 s. It was trained to reproduce each HuBERT frame five frames (100 ms) later, so its features describe the speech with a delay of about 100 ms and a detector built on it reacts that much after the word ends. For streaming, keep 40,000 samples (2.5 s) of context so streamed frames equal offline ones. The log-mel front end is part of the graph. wakehubert_int8.onnx is a static int8 version, with the front end kept in float, whose features agree with float at a mean cosine of 0.998.

import numpy as np, onnxruntime as ort
from huggingface_hub import hf_hub_download
path = hf_hub_download("TigreGotico/wakehubert-tiny", "wakehubert.onnx")
sess = ort.InferenceSession(path)
feats = sess.run(None, {"waveform": np.zeros((1, 24000), np.float32)})[0]  # (1, 75, 128)

In wakeforge: ww_trainer-train --featurizer wakehubert ... downloads and uses it.

How it was made

The student is a fixed causal log-mel front end (64 bins), a strided convolution to 50 frames per second, eight dilated depthwise-separable convolution blocks with 256 channels, and a 1×1 projection to 128 features: convolution, batch norm and ReLU only, which quantises well. It was trained for 30,000 steps to predict standardised HuBERT-base layers 4, 8 and 12 (L1 plus log-sigmoid cosine, as in DistilHuBERT), with the teacher hearing clean speech while the student heard it with AudioSet, MUSAN and non-speech noise, room reverberation and one to three background talkers never louder than the voice. A quarter of the training items were non-speech sounds heard identically by both. Spans of the student's input log-mel were masked during training (probability 0.065 per frame, 10-frame spans), which was the largest single gain in robustness found in the experiments.

Training speech: LibriSpeech (960 h), Multilingual LibriSpeech (seven languages) and a language-balanced sample of Multilingual Spoken Words (41 languages), cut into 600,000 two-second crops.

Results

Single runs. A GRU classifier trained only on 900 synthetic "alexa" clips (TTS voices cloned with voice conversion, with noise, babble, reverberation, speed and gain augmentation) was scored on the real speakers of the Picovoice wake-word benchmark (315 recordings), with the threshold chosen on separate calibration audio (LibriSpeech dev-clean and babble made from it) for 0.5 false activations per hour, and false activations then measured on 6.5 h of held-out streams (LibriSpeech test-clean, three-talker test-other babble, held-out non-speech):

Extractor Recall, quiet (95% interval) Recall in babble at 10 / 5 / 0 dB False activations per hour measured
WakeHuBERT tiny (0.64M) 95% (92–97) 94 / 85 / 56% 0.31
Same architecture without masking 92% (89–95) 90 / 80 / 39% 0.46

The recall intervals come from resampling the 315 recordings; differences of a few points between extractors are within noise, and two classifier types on the same extractor can differ by several points.

PyTorch weights

model.safetensors holds the float32 weights of the student (641,792 trained parameters, plus 182,514 fixed values: 173,664 in the log-mel front end and 8,850 batch-norm statistics) and student.py defines the network in plain PyTorch with no dependency beyond torch and safetensors. The weights are the checkpoint that wakehubert.onnx was exported from, so they can be fine-tuned or exported to other formats. The three linear heads that predicted the HuBERT layers during distillation are not included.

from huggingface_hub import snapshot_download
import sys, torch; d = snapshot_download("TigreGotico/wakehubert-tiny"); sys.path.insert(0, d)
from student import load_wakehubert_tiny
model = load_wakehubert_tiny(f"{d}/model.safetensors")
feats = model(torch.zeros(1, 24000))  # (1, 75, 128)
torch.onnx.export(model, torch.zeros(1, 16000), "wakehubert.onnx", input_names=["waveform"], output_names=["features"], dynamic_axes={"waveform": {0: "batch", 1: "samples"}, "features": {0: "batch", 1: "frames"}}, opset_version=17, dynamo=False)

The PyTorch model and wakehubert.onnx agree to within 1e-5 on the 128-dimensional features (largest absolute difference 9.5e-6 over five random three-second waveforms and one ten-second LibriSpeech clip), and exporting the PyTorch model again gives an ONNX graph whose outputs match the published one to within 5e-6. The distillation and training code is in wakeforge.

Full training checkpoints

training/best.pt (step 26,000, the checkpoint wakehubert.onnx and model.safetensors come from) and training/last.pt (step 30,000) are the complete PyTorch checkpoints of the distillation run, for continuing the distillation rather than only fine-tuning the student. Each is a dict with student (the student state_dict), norm (the standardisation statistics of the HuBERT targets), opt and sched (optimizer and learning-rate scheduler state), step, best and meta; the three linear heads that predicted HuBERT layers 4, 8 and 12 are inside student under proj.*. They are pickle files: load them with torch.load(path, map_location="cpu", weights_only=False) only if you trust this repository. The distillation script that reads them is in wakeforge.

Limitations

The evaluation covers one English wake word from one public benchmark; results for other words, languages and devices are not measured here. Agreement with the teacher is a poor predictor of detection quality, so the extractor should be judged by a detector trained on it.

Licence and attribution

Released under the Apache License 2.0, the licence of its teacher (HuBERT-base, Apache 2.0). Trained with LibriSpeech, Multilingual LibriSpeech and Multilingual Spoken Words (all CC BY 4.0), MUSAN, AudioSet-derived noise and room impulse responses, and distilled from facebook/hubert-base-ls960.

Downloads last month
834
Safetensors
Model size
824k params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TigreGotico/wakehubert-tiny

Quantized
(10)
this model
Quantizations
3 models

Datasets used to train TigreGotico/wakehubert-tiny

Space using TigreGotico/wakehubert-tiny 1

Collection including TigreGotico/wakehubert-tiny

Evaluation results

  • Recall, quiet, at 0.5 false activations per hour on Picovoice wake-word benchmark, alexa, real speakers (315 recordings)
    self-reported
    0.950
  • Recall in 5 dB babble on Picovoice wake-word benchmark, alexa, real speakers (315 recordings)
    self-reported
    0.850
  • False activations per hour on 6.5 h of held-out streams on Picovoice wake-word benchmark, alexa, real speakers (315 recordings)
    self-reported
    0.310