【Short Technical Report & Usage】NexteraBERT: Input-Dependent Gating Liquid Mixer and Length-Adaptive Attention for Fast, Long-Context Bidirectional Encoders
Rikka Botan
Independent Researcher, Japan
https://rikka-botan.github.io
Abstract
- Models: 130B tokens (main model) · 13B tokens · 1.3B tokens
- Paper: NexteraBERT.pdf
- Code: Rikka-Botan/NexteraBERT (MIT license)
- Pretraining data: RikkaBotan/NexteraBERT-data-mix
1 Introduction
Encoders such as BERT read the whole input in one forward pass, which makes them the standard choice for classification, semantic similarity, and retrieval, including the retrieval step of RAG. ModernBERT brought many advances from decoder models to encoders: rotary positions, alternating local and global attention, efficient kernels, an 8,192-token context, and 2T pretraining tokens. OptiBERT measured how much quality a given pretraining budget gives a standard Transformer encoder.
Three problems remain:
- Long inputs are expensive.
- Quality drops beyond the training length.
- Pretraining is expensive.
NexteraBERT addresses the three together by an architecture and a training recipe.
2 Method(Architecture)
Every block has a token mixer and a squared-ReLU MLP, each with pre-normalization and a residual connection. What changes is the token mixer.
| Token mixer | Blocks | What it does | Cost in input length |
|---|---|---|---|
| SnowLily | 8 | local mixing with an input-dependent gated convolution | linear |
| Sliding-window attention | 5 | attention within a 256-token window | linear |
| Full attention with SSSMax | 3 | global attention with Scalable Softmax and Gated Attention | quadratic |
| HRA | 2 | attention between mean-pooled bands of 4 tokens | quadratic, 16× fewer pairs |
Thirteen of the 18 blocks are local and cost time linear in the input length, and only five blocks connect distant tokens.
- Blocks 1–6 (surface and phrase-level features): four SnowLily and two window-attention blocks, with no global attention.
- Blocks 7–12 (syntactic features): full attention at blocks 7 and 12, with local mixers between them.
- Blocks 13–18 (semantic features and long-distance dependencies): the most global computation, with HRA at blocks 13 and 17 and full attention at block 16.
The layer roles come from studies of BERT and serve as a design guideline.
3 Method(Training recipe)
The data mix is 90.25% FineWeb-Edu, 4.75% DCLM, and 5% StarCoderData. Pretraining uses standard MLM (80/10/10 replacement), AdamW, and bf16, in two phases:
| Phase 1 | Phase 2 | |
|---|---|---|
| Tokens (130B model) | 120B | 10B |
| Input length | up to 1,024 | up to 8,192; 1,024-token batches on 40% of steps |
| Masking rate | 40% → 20%, linear | 20% |
| Peak learning rate | 8e-4 | 8e-5 |
4 Results
(a, b) GLUE and MTEB v2 against pretraining tokens: NexteraBERT at 1.3B, 13B, and 130B tokens, ModernBERT-base (2T) and NeoBERT (2.1T) evaluated under the same protocol, and published OptiBERTneo values. (c) Throughput on an H200 NVL with 65,536 tokens per batch. (d) Exponentiated masked-token loss; the dotted line marks 8,192 tokens, the longest training length.
4.1 Downstream quality
Both models below are fine-tuned and evaluated with the same scripts, which are in the GitHub repository. Better scores are in bold.
| Benchmark | Metric | NexteraBERT | ModernBERT-base |
|---|---|---|---|
| Pretraining tokens | tokens | 130B | 2T |
| GLUE (8 tasks) | mean | 87.90 | 87.97 |
| MTEB v2 (English, 41 tasks) | mean over task types | 54.70 | 53.63 |
| BEIR (15 datasets) | nDCG@10 | 43.89 | 43.09 |
| NanoBEIR (13 subsets) | nDCG@10 | 55.92 | 53.76 |
| Code retrieval (CodeSearchNet, StackOverflowQA) | nDCG@10, mean | 61.46 | 62.92 |
| MultiLongDocRetrieval, 8,192 tokens | nDCG@10 | 36.25 | 31.19 |
4.2 Training efficiency
| Pretraining tokens | NexteraBERT | OptiBERTneo | Gain |
|---|---|---|---|
| 1.3B | 47.82 | 47.40 | +0.42 |
| 13B | 51.92 | 50.50 | +1.42 |
| 130B | 54.70 | 52.60 | +2.10 |
MTEB v2, mean over task types.
NexteraBERT is ahead at every budget with 30% less estimated compute. Across budgets, the best compute-optimal OptiBERT model reaches 51.60 at 7.0 × 10¹⁹ FLOPs; our 13B-token model reaches 51.92 at 1.33 × 10¹⁹ FLOPs, 5.27 times less.
4.3 Long inputs: speed and extrapolation
We timed on one H200 NVL (bf16, torch.compile) with 65,536 tokens in every batch, so the batch size falls from 64 at 1k tokens to 1 at 64k.
| Batch latency (ms) | 1k | 8k | 64k |
|---|---|---|---|
| NexteraBERT | 70.7 | 100.2 | 276.5 |
| ModernBERT-base | 59.4 | 239.3 | 1,442.4 |
The advantage over ModernBERT-base grows with length: 2.39 times at 8k, 3.57 times at 16k, and 5.22 times at 64k.
For extrapolation, we measure the exponentiated masked-token loss on FineWeb-Edu windows of 1,024 to 65,536 tokens (15% masking, 25 draws).
| Exponentiated MLM loss | 8,192 tokens | 65,536 tokens | Increase |
|---|---|---|---|
| NexteraBERT | 3.004 | 3.118 | +3.8% |
| ModernBERT-base | 3.265 | 5.977 | +83.0% |
5 How to use NexteraBERT
5.1 With Transformers
The Hub repositories include the model code, so pass trust_remote_code=True. The snippets below were tested with transformers 5.12 and PyTorch 2.12.
pip install -U transformers torch
Masked-token prediction
from transformers import AutoModelForMaskedLM, AutoTokenizer, pipeline
repo = "RikkaBotan/NexteraBERT-Mezzoforte-220M-en"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True)
fill = pipeline("fill-mask", model=model, tokenizer=tokenizer)
for p in fill("The capital of France is [MASK].", top_k=3):
print(p["token_str"], round(p["score"], 3))
# Paris 0.837
# Lyon 0.057
# Nice 0.019
Token features and a mean-pooled embedding
import torch
from transformers import AutoModel, AutoTokenizer
repo = "RikkaBotan/NexteraBERT-Mezzoforte-220M-en"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModel.from_pretrained(repo, trust_remote_code=True).eval()
batch = tokenizer(["hello world", "a second sentence"], padding=True, return_tensors="pt")
with torch.no_grad():
hidden = model(**batch).last_hidden_state # (batch, length, 1024)
mask = batch["attention_mask"].unsqueeze(-1).to(hidden.dtype)
embedding = (hidden * mask).sum(1) / mask.sum(1) # mean pooling
NexteraBERT is a pretrained encoder, not a sentence-embedding model. Fine-tune it (for example with contrastive learning, as for the MTEB and retrieval results above) before using it for similarity or retrieval.
Batches of mixed lengths
model = AutoModel.from_pretrained(repo, trust_remote_code=True, unpadding=True)
With unpadding=True, the encoder packs the real tokens of all rows into one stream, so no computation is spent on padding, and each row is encoded exactly as if it were alone in the batch.
Classification
from transformers import AutoModelForSequenceClassification
model = AutoModelForSequenceClassification.from_pretrained(
"RikkaBotan/NexteraBERT-Mezzoforte-220M-en", trust_remote_code=True, num_labels=3
)
The classification head is newly initialized, so train it on your task first, for example with the Trainer.
Simple finetune with sst2
import numpy as np
import torch
from datasets import load_dataset
from transformers import (
AutoModelForSequenceClassification,
AutoTokenizer,
DataCollatorWithPadding,
Trainer,
TrainingArguments,
set_seed,
)
set_seed(42)
model_name = "RikkaBotan/NexteraBERT-Mezzoforte-220M-en"
dataset = load_dataset("nyu-mll/glue", "sst2")
train_dataset = dataset["train"].shuffle(seed=42).select(range(10000))
eval_dataset = dataset["validation"]
tokenizer = AutoTokenizer.from_pretrained(model_name)
def tokenize(examples):
return tokenizer(examples["sentence"], truncation=True, max_length=128)
train_dataset = train_dataset.map(tokenize, batched=True)
eval_dataset = eval_dataset.map(tokenize, batched=True)
id2label = {0: "negative", 1: "positive"}
label2id = {"negative": 0, "positive": 1}
model = AutoModelForSequenceClassification.from_pretrained(
model_name,
trust_remote_code=True,
num_labels=2,
id2label=id2label,
label2id=label2id,
)
torch.nn.init.normal_(model.classifier.weight, mean=0.0, std=0.02)
torch.nn.init.zeros_(model.classifier.bias)
def compute_metrics(eval_pred):
logits, labels = eval_pred
predictions = np.argmax(logits, axis=-1)
return {"accuracy": (predictions == labels).mean()}
training_args = TrainingArguments(
output_dir="nexterabert-sst2",
learning_rate=4e-5,
per_device_train_batch_size=16,
per_device_eval_batch_size=64,
num_train_epochs=1,
weight_decay=0.01,
warmup_steps=0.06,
eval_strategy="epoch",
save_strategy="no",
logging_steps=100,
report_to="none",
seed=42,
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=train_dataset,
eval_dataset=eval_dataset,
processing_class=tokenizer,
data_collator=DataCollatorWithPadding(tokenizer),
compute_metrics=compute_metrics,
)
trainer.train()
print(trainer.evaluate())
trainer.save_model("nexterabert-sst2/final")
5.2 Pretrain and Evaluation with the GitHub repository
Rikka-Botan/NexteraBERT contains the model code and everything used for the paper: data preparation, the two-phase pretraining, and all evaluations. The project is managed with uv:
git clone https://github.com/Rikka-Botan/NexteraBERT
cd NexteraBERT
uv sync --extra train --extra eval
source .venv/bin/activate
If the default PyTorch wheel of your platform has no CUDA support, install the matching CUDA build from pytorch.org.
Pretraining. One command runs the whole pipeline: it downloads the tokenized data from the Hub, trains both phases, exports the model, and runs GLUE and the SimCSE + MTEB evaluation. If HF_REPO_ID is set, it also uploads the model.
NPROC=1 bash scripts/run_pipeline_bert.sh
python scripts/pretrain.py --config configs/pretrain_bert_phase1.yaml # up to 1,024 tokens
python scripts/pretrain.py --config configs/pretrain_bert_phase2.yaml # up to 8,192 tokens, starts from phase 1
Metrics go to Weights & Biases (run wandb login, or set WANDB_MODE=offline to keep the logs local).
Evaluation.
| Result | Command |
|---|---|
| GLUE | bash scripts/eval_nlu.sh |
| MTEB v2 (SimCSE fine-tuning, then 41 tasks) | bash scripts/eval_mteb.sh |
| BEIR (15 datasets) and MultiLongDocRetrieval | bash scripts/eval_dpr.sh |
| CodeSearchNet and StackOverflowQA | bash scripts/eval_code.sh |
| NanoBEIR | bash scripts/eval_nanobeir.sh |
| Throughput vs. input length | python src/nexterabert/model_speedbench.py |
| Masked-token loss vs. input length | python src/nexterabert/model_pplbench.py |
Other encoders go through the same stages, so a baseline can be evaluated under the identical protocol:
MODEL=answerdotai/ModernBERT-base bash scripts/eval_nlu.sh
6 Conclusion
With about one fifteenth of the pretraining tokens, NexteraBERT matches ModernBERT-base on GLUE and exceeds it on MTEB v2. It is several times faster on long inputs, and its maskedtoken loss stays stable up to eight times its longest training length. Even training with just 130B can achieve downstream task performance equal to or better than ModernBERT, so I think even small labs can conduct sufficient follow-up research.
Acknowledgements
My interest in this topic originated from reading the paper of ModernBERT, so I especially want to thank the authors, Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Nathan Cooper, Griffin Adams, Jeremy Howard, Iacopo Poli .
I thank the developers of transformers, python and pytorch.
I thank all the researchers for their efforts to date.
I thank Japan's high standard of education.
And most of all, thank you for your interest in this blog.
About us
Japanese independent researcher having shy and pampered personality. Twin-tail hair is a charm point. Interested in nlp. Usually using python and C.
Please contact us if you have any requests for joint research, writing, speaking engagements, or employment.
Reference
Alexander Amini, Anna Banaszak, Harold Benoit et al. LFM2 Technical Report. arXiv:2511.23404, 2025.
Ken M. Nakanishi. Scalable-Softmax Is Superior for Attention. arXiv:2501.19399, 2025.
Raymond Li, Loubna Ben Allal, Yangtian Zi et al. StarCoder: may the source be with you! TMLR 2023.
Liquid AI. LFM2.5-Encoders: Fast at Long Context, Even on CPU. Liquid AI blog, 2026.


