CallEnhancer

CallEnhancer restores call-centre speech. It takes narrowband, codec'd, noisy phone audio, such as 8 kHz G.711 or GSM calls, and returns clean speech at 48 kHz.

This is version 2. Version 1 stays available at the v1 tag.

What changed in v2

  • The full band. v2's high-band distance to clean speech is 5.5 dB on our 1,289 call-centre clips, against 6.9 dB for v1, and 3.9 dB against 5.3 dB on 58 real calls. Clean studio recordings score 2.0 to 2.3 dB.
  • Intelligibility. On the 58 real calls v2 reads 30.8 % nCER and v1 30.1 %. The mean per-call difference is +1.1 points (95 % CI [-0.9, +3.5]), within noise. v2 is better on 27 calls and worse on 30.
  • Deletions. v2 deletes less speech. It mutes 10.9 % of speech against 13.2 % for v1, at 40.5 % against 41.9 % silence filled: more than a change in loudness balance explains.
  • Two parties in one channel. v2 keeps a quieter second voice. With the two sides of our 58 real calls mixed into one channel, v2 dropped 4.4 % of the quieter party's speech and v1 4.1 %. v2 still drops more than a quarter of it in 4 of 56 calls, v1 in 1.
  • Naturalness. About the same. SOTA-MOS is 2.77 for v2 and 2.70 for v1 (300 clips). Two runs of one recipe differ by 0.07 to 0.08.
  • Same size, same files. A w2v-BERT 2.0 encoder with a LoRA adapter and a 188M DAC decoder. The weights keep v1's paths, so existing download code keeps working.

Comparison

We ran CallEnhancer v2, CallEnhancer v1 and three open restoration models on the same call-centre audio, through the same scripts.

Median spectrum on 1,289 call-centre clips

Each curve is a system's median spectrum, relative to its own 1 to 4 kHz band, on speech-active frames. Black is clean speech. The number in the legend is the high-band distance (HBD): the RMS gap to the black curve from 4 to 20 kHz. Lower is closer to clean speech.

model output HBD, 1,289 clips (dB) โ†“ HBD, 58 real calls (dB) โ†“ nCER, 58 real calls (%) โ†“ speech muted (%) โ†“ silence filled (%) SOTA-MOS โ†‘ DNSMOS overall โ†‘
CallEnhancer v2 (this release) 48 kHz 5.5 3.9 30.8 10.9 40.5 2.77 2.97
CallEnhancer v1 48 kHz 6.9 5.3 30.1 13.2 41.9 2.70 2.76
sidon-v0.1 48 kHz 4.4 5.2 30.8 28.2 20.9 2.64 3.17
VoiceFixer 44.1 kHz 3.0 2.1 50.0 27.1 52.9 2.19 2.78
Resemble Enhance 44.1 kHz 4.8 8.6 48.8 19.5 15.7 2.28 3.05
input (telephone audio) 8 to 16 kHz 51.3 49.2 48.8

Bold marks the best system in each column. Silence filled has no best: it is the counter-metric to speech muted, not a score (see How we measured). The input row is for reference: unprocessed calls read 48.8 % nCER.

Median spectrum on 58 real calls

Character error rate on 58 real calls

Speech muted against silence filled

SOTA-MOS on 300 clips

How we measured

  • Test audio. 1,289 in-domain call-centre clips (English, Malay and Mandarin, 9.4 hours) and 58 real calls with human transcripts (116 legs, one party per channel). Both are customer data and stay private. We publish aggregates only.
  • Same input for everyone. Every system restores the same files. VoiceFixer and Resemble Enhance output 44.1 kHz; we resample them to 48 kHz before scoring.
  • Settings. CallEnhancer v1 and v2: infer_callcentre.py defaults (4-second windows, native 48 kHz output). sidon-v0.1: the released TorchScript model. VoiceFixer: restore, mode 0. Resemble Enhance: enhance with its command-line defaults (64 function evaluations, midpoint solver, denoise strength 1.0, temperature 0.5).
  • HBD. The RMS gap from 4 to 20 kHz between a system's median spectrum and clean speech, both relative to the 1 to 4 kHz band.
  • nCER. Character error rate of mesolitica/Malaysian-whisper-large-v3-turbo-v3 against the human transcripts, after LLM normalisation, mean of 3 repeats. The two parties are restored separately and mixed at matched levels. All systems were scored in one session.
  • Speech muted. The share of speech (20 ms frames) that the output drops by more than 6 dB, on the clips whose reference is speech. Each signal is measured against its own 90th-percentile frame, so a gain change does not count. A plain band-pass that deletes nothing scores 5.9 %.
  • Silence filled. The counter-metric to speech muted. It is the share of frames that are silent in the input (40 dB or more below its 90th-percentile frame) that the output raises by more than 6 dB. It is not a score. High means energy added to the pauses. Low is no merit either: an output that pushes its quiet parts down fills little silence and mutes more speech. A change in loudness balance alone moves the two in opposite directions. In our fault-injection test, compressing an output's dynamic range by 10 % lowered muting by 1.6 points and raised silence filled by 18; a real deletion raises muting with silence filled flat. So compare muting at similar silence filled. The shaded band in the deletions figure is where v1 would sit if only its loudness balance changed.
  • SOTA-MOS. Scicom-intl/HighRateMOS-VoiceMOS2025 on 300 paired clips at 48 kHz.
  • One-channel mixes. The two sides of each of the 58 real calls averaged into one channel at their recorded levels (the sides' speech levels differ by 11.4 dB at the median) and restored as one file. A party's frame counts when its side is within 25 dB of that side's 90th-percentile frame and 15 dB or more above the other side. Muted as for speech muted, with the outputs resampled to 8 kHz so the input and the output share the band.
  • DNSMOS P.835 overall, the median over the clips it can score. It cannot score the 6 clips under 1 second, or an output whose middle 10 seconds are silent. The medians cover 1,283 clips (sidon-v0.1 1,282, Resemble Enhance 1,243).
  • Silent outputs. Clips whose whole output stays under -60 dBFS: CallEnhancer v2 0, CallEnhancer v1 0, sidon-v0.1 0, VoiceFixer 0, Resemble Enhance 41. 7 of the 1,289 inputs are silent. Where the input has speech, the speech-muted column counts a silent output as a deletion.

Quick start

pip install torch torchaudio "transformers>=4.56" "descript-audio-codec>=1.0.0" soundfile "huggingface_hub>=0.34"

hf download Scicom-intl/CallEnhancer infer_callcentre.py expand_decoder.py \
    fe_callcentre/fe_adapter_full.pt decoder_callcentre/decoder_only.pt --local-dir CallEnhancer
cd CallEnhancer

python infer_callcentre.py --input your_call.wav --out-dir out \
    --fe-adapter fe_callcentre/fe_adapter_full.pt \
    --decoder decoder_callcentre/decoder_only.pt --device cuda
  • out/your_call_restored48k.wav is the restored speech at 48 kHz.
  • out/your_call_orig48k.wav is the input, upsampled with no model, for an A/B listen.
  • --input takes a file or a directory (.wav, .flac, .mp3, .ogg, .opus, .m4a). Stereo calls are restored per channel. One channel holding both parties works too: v2 keeps the quieter voice.
  • The default restores in 4-second windows with a 2-second crossfade. Keep it: the encoder's feature normalisation depends on clip length, and a single pass over a long file drops speech. Use --device cpu without a GPU.
  • v1 stays at the tag v1. To run it, keep this script and fetch only v1's weights: hf download Scicom-intl/CallEnhancer fe_callcentre/fe_adapter_full.pt decoder_callcentre/decoder_only.pt --revision v1 --local-dir v1, then pass them to --fe-adapter and --decoder. v1's own script restores in a single pass by default.

Python

import numpy as np, soundfile as sf, torch, torchaudio
from huggingface_hub import hf_hub_download
from transformers import AutoFeatureExtractor
import infer_callcentre as ic          # from the download above, next to expand_decoder.py

torch.set_float32_matmul_precision("medium")                 # as the command line sets it: same output, bit for bit
dev = torch.device("cuda" if torch.cuda.is_available() else "cpu")
fe = ic.load_fe(hf_hub_download("Scicom-intl/CallEnhancer", "fe_callcentre/fe_adapter_full.pt"), dev)
dec, out_sr = ic.load_decoder(hf_hub_download("Scicom-intl/CallEnhancer", "decoder_callcentre/decoder_only.pt"), dev)
proc = AutoFeatureExtractor.from_pretrained("facebook/w2v-bert-2.0")

x, sr = sf.read("your_call.wav", dtype="float32", always_2d=True)
x16 = torchaudio.functional.resample(torch.from_numpy(x[:, 0])[None], sr, 16000)[0].numpy()   # one channel
x16 = x16 / (np.abs(x16).max() + 1e-9) * 0.95
y = ic.restore_channel(x16, fe, dec, proc, dev, chunk_s=4.0, bf16=True, out_sr=out_sr)
y = y / (np.abs(y).max() + 1e-9) * 0.97                       # peak level, as the command line writes it
sf.write("restored48k.wav", y, out_sr)

Files

path what it is
fe_callcentre/fe_adapter_full.pt encoder adapter for inference: 96 LoRA tensors and 48 output_dense biases on w2v-BERT 2.0
decoder_callcentre/decoder_only.pt decoder for inference: the 188M DAC decoder, 48 kHz
decoder_callcentre/last.pt the full training state at step 150,000 (decoder, adapter, discriminator, optimisers), to continue training
infer_callcentre.py, expand_decoder.py the inference script and the decoder builder it imports
examples/ five public CallHome calls (TalkBank), each one channel with both parties: the input, and the v2 and v1 restorations
assets/ the figures on this card

Limitations

  • The benchmarks are one operator's calls in English, Malay and Mandarin. Other lines and languages may differ.
  • Sidon-style restoration resynthesises speech. Above 8 kHz the output is generated from the encoder's features, not recovered.
  • SOTA-MOS mostly judges the band below 8 kHz. Use HBD and the spectra for the band above it.
  • The test audio is customer data. We publish aggregates only, so the numbers cannot be reproduced outside Scicom.

License

CC BY-NC 4.0. Some training noise comes from non-commercial sources.

Built by Scicom (MSC) Berhad.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Scicom-intl/CallEnhancer

Adapter
(2)
this model