【Short Technical Report & Usage】NexteraBERT: Input-Dependent Gating Liquid Mixer and Length-Adaptive Attention for Fast, Long-Context Bidirectional Encoders

Community Article
Published September 28, 2026

Rikka Botan
Independent Researcher, Japan
https://rikka-botan.github.io

Abstract

NexteraBERT is a 212.6M-parameter bidirectional encoder pretrained with masked language modeling (MLM). Instead of repeating one or two block types, it arranges four token mixers. With 130B pretraining tokens, about 15 times fewer than ModernBERT, it matches ModernBERT-base on GLUE (87.90 vs. 87.97) and exceeds it on MTEB v2 (54.70 vs. 53.63), with both models fine-tuned under the same protocol. It reaches 5.22 times the throughput of ModernBERT-base at 65,536 tokens on an NVIDIA H200 NVL and its masked-token loss rises by only 3.8% from 8,192 to 65,536 tokens, eight times its longest training length, against 83.0% for ModernBERT-base.

NexteraBERT logo

1 Introduction

Encoders such as BERT read the whole input in one forward pass, which makes them the standard choice for classification, semantic similarity, and retrieval, including the retrieval step of RAG. ModernBERT brought many advances from decoder models to encoders: rotary positions, alternating local and global attention, efficient kernels, an 8,192-token context, and 2T pretraining tokens. OptiBERT measured how much quality a given pretraining budget gives a standard Transformer encoder.

Three problems remain:

  1. Long inputs are expensive.
  2. Quality drops beyond the training length.
  3. Pretraining is expensive.

NexteraBERT addresses the three together by an architecture and a training recipe.

2 Method(Architecture)

NexteraBERT architecture

Every block has a token mixer and a squared-ReLU MLP, each with pre-normalization and a residual connection. What changes is the token mixer.

Token mixer Blocks What it does Cost in input length
SnowLily 8 local mixing with an input-dependent gated convolution linear
Sliding-window attention 5 attention within a 256-token window linear
Full attention with SSSMax 3 global attention with Scalable Softmax and Gated Attention quadratic
HRA 2 attention between mean-pooled bands of 4 tokens quadratic, 16× fewer pairs

Thirteen of the 18 blocks are local and cost time linear in the input length, and only five blocks connect distant tokens.

  • Blocks 1–6 (surface and phrase-level features): four SnowLily and two window-attention blocks, with no global attention.
  • Blocks 7–12 (syntactic features): full attention at blocks 7 and 12, with local mixers between them.
  • Blocks 13–18 (semantic features and long-distance dependencies): the most global computation, with HRA at blocks 13 and 17 and full attention at block 16.

The layer roles come from studies of BERT and serve as a design guideline.

3 Method(Training recipe)

The data mix is 90.25% FineWeb-Edu, 4.75% DCLM, and 5% StarCoderData. Pretraining uses standard MLM (80/10/10 replacement), AdamW, and bf16, in two phases:

Phase 1 Phase 2
Tokens (130B model) 120B 10B
Input length up to 1,024 up to 8,192; 1,024-token batches on 40% of steps
Masking rate 40% → 20%, linear 20%
Peak learning rate 8e-4 8e-5

4 Results

Main results of NexteraBERT

(a, b) GLUE and MTEB v2 against pretraining tokens: NexteraBERT at 1.3B, 13B, and 130B tokens, ModernBERT-base (2T) and NeoBERT (2.1T) evaluated under the same protocol, and published OptiBERTneo values. (c) Throughput on an H200 NVL with 65,536 tokens per batch. (d) Exponentiated masked-token loss; the dotted line marks 8,192 tokens, the longest training length.

4.1 Downstream quality

Both models below are fine-tuned and evaluated with the same scripts, which are in the GitHub repository. Better scores are in bold.

Benchmark Metric NexteraBERT ModernBERT-base
Pretraining tokens tokens 130B 2T
GLUE (8 tasks) mean 87.90 87.97
MTEB v2 (English, 41 tasks) mean over task types 54.70 53.63
BEIR (15 datasets) nDCG@10 43.89 43.09
NanoBEIR (13 subsets) nDCG@10 55.92 53.76
Code retrieval (CodeSearchNet, StackOverflowQA) nDCG@10, mean 61.46 62.92
MultiLongDocRetrieval, 8,192 tokens nDCG@10 36.25 31.19

4.2 Training efficiency

Pretraining tokens NexteraBERT OptiBERTneo Gain
1.3B 47.82 47.40 +0.42
13B 51.92 50.50 +1.42
130B 54.70 52.60 +2.10

MTEB v2, mean over task types.

NexteraBERT is ahead at every budget with 30% less estimated compute. Across budgets, the best compute-optimal OptiBERT model reaches 51.60 at 7.0 × 10¹⁹ FLOPs; our 13B-token model reaches 51.92 at 1.33 × 10¹⁹ FLOPs, 5.27 times less.

4.3 Long inputs: speed and extrapolation

We timed on one H200 NVL (bf16, torch.compile) with 65,536 tokens in every batch, so the batch size falls from 64 at 1k tokens to 1 at 64k.

Batch latency (ms) 1k 8k 64k
NexteraBERT 70.7 100.2 276.5
ModernBERT-base 59.4 239.3 1,442.4

The advantage over ModernBERT-base grows with length: 2.39 times at 8k, 3.57 times at 16k, and 5.22 times at 64k.

For extrapolation, we measure the exponentiated masked-token loss on FineWeb-Edu windows of 1,024 to 65,536 tokens (15% masking, 25 draws).

Exponentiated MLM loss 8,192 tokens 65,536 tokens Increase
NexteraBERT 3.004 3.118 +3.8%
ModernBERT-base 3.265 5.977 +83.0%

5 How to use NexteraBERT

5.1 With Transformers

The Hub repositories include the model code, so pass trust_remote_code=True. The snippets below were tested with transformers 5.12 and PyTorch 2.12.

pip install -U transformers torch

Masked-token prediction

from transformers import AutoModelForMaskedLM, AutoTokenizer, pipeline

repo = "RikkaBotan/NexteraBERT-Mezzoforte-220M-en"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForMaskedLM.from_pretrained(repo, trust_remote_code=True)

fill = pipeline("fill-mask", model=model, tokenizer=tokenizer)
for p in fill("The capital of France is [MASK].", top_k=3):
    print(p["token_str"], round(p["score"], 3))
# Paris 0.837
# Lyon 0.057
# Nice 0.019

Token features and a mean-pooled embedding

import torch
from transformers import AutoModel, AutoTokenizer

repo = "RikkaBotan/NexteraBERT-Mezzoforte-220M-en"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModel.from_pretrained(repo, trust_remote_code=True).eval()

batch = tokenizer(["hello world", "a second sentence"], padding=True, return_tensors="pt")
with torch.no_grad():
    hidden = model(**batch).last_hidden_state          # (batch, length, 1024)

mask = batch["attention_mask"].unsqueeze(-1).to(hidden.dtype)
embedding = (hidden * mask).sum(1) / mask.sum(1)       # mean pooling

NexteraBERT is a pretrained encoder, not a sentence-embedding model. Fine-tune it (for example with contrastive learning, as for the MTEB and retrieval results above) before using it for similarity or retrieval.

Batches of mixed lengths

model = AutoModel.from_pretrained(repo, trust_remote_code=True, unpadding=True)

With unpadding=True, the encoder packs the real tokens of all rows into one stream, so no computation is spent on padding, and each row is encoded exactly as if it were alone in the batch.

Classification

from transformers import AutoModelForSequenceClassification

model = AutoModelForSequenceClassification.from_pretrained(
    "RikkaBotan/NexteraBERT-Mezzoforte-220M-en", trust_remote_code=True, num_labels=3
)

The classification head is newly initialized, so train it on your task first, for example with the Trainer.

Simple finetune with sst2

import numpy as np
import torch
from datasets import load_dataset
from transformers import (
    AutoModelForSequenceClassification,
    AutoTokenizer,
    DataCollatorWithPadding,
    Trainer,
    TrainingArguments,
    set_seed,
)

set_seed(42)
model_name = "RikkaBotan/NexteraBERT-Mezzoforte-220M-en"
dataset = load_dataset("nyu-mll/glue", "sst2")
train_dataset = dataset["train"].shuffle(seed=42).select(range(10000))
eval_dataset = dataset["validation"]

tokenizer = AutoTokenizer.from_pretrained(model_name)
def tokenize(examples):
    return tokenizer(examples["sentence"], truncation=True, max_length=128)

train_dataset = train_dataset.map(tokenize, batched=True)
eval_dataset = eval_dataset.map(tokenize, batched=True)

id2label = {0: "negative", 1: "positive"}
label2id = {"negative": 0, "positive": 1}

model = AutoModelForSequenceClassification.from_pretrained(
    model_name,
    trust_remote_code=True,
    num_labels=2,
    id2label=id2label,
    label2id=label2id,
)

torch.nn.init.normal_(model.classifier.weight, mean=0.0, std=0.02)
torch.nn.init.zeros_(model.classifier.bias)

def compute_metrics(eval_pred):
    logits, labels = eval_pred
    predictions = np.argmax(logits, axis=-1)
    return {"accuracy": (predictions == labels).mean()}

training_args = TrainingArguments(
    output_dir="nexterabert-sst2",
    learning_rate=4e-5, 
    per_device_train_batch_size=16,
    per_device_eval_batch_size=64,
    num_train_epochs=1,
    weight_decay=0.01,
    warmup_steps=0.06, 
    eval_strategy="epoch",
    save_strategy="no", 
    logging_steps=100, 
    report_to="none", 
    seed=42,
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=train_dataset,
    eval_dataset=eval_dataset,
    processing_class=tokenizer, 
    data_collator=DataCollatorWithPadding(tokenizer), 
    compute_metrics=compute_metrics,
)

trainer.train()
print(trainer.evaluate())
trainer.save_model("nexterabert-sst2/final") 

5.2 Pretrain and Evaluation with the GitHub repository

Rikka-Botan/NexteraBERT contains the model code and everything used for the paper: data preparation, the two-phase pretraining, and all evaluations. The project is managed with uv:

git clone https://github.com/Rikka-Botan/NexteraBERT
cd NexteraBERT
uv sync --extra train --extra eval
source .venv/bin/activate

If the default PyTorch wheel of your platform has no CUDA support, install the matching CUDA build from pytorch.org.

Pretraining. One command runs the whole pipeline: it downloads the tokenized data from the Hub, trains both phases, exports the model, and runs GLUE and the SimCSE + MTEB evaluation. If HF_REPO_ID is set, it also uploads the model.

NPROC=1 bash scripts/run_pipeline_bert.sh
python scripts/pretrain.py --config configs/pretrain_bert_phase1.yaml   # up to 1,024 tokens
python scripts/pretrain.py --config configs/pretrain_bert_phase2.yaml   # up to 8,192 tokens, starts from phase 1

Metrics go to Weights & Biases (run wandb login, or set WANDB_MODE=offline to keep the logs local).

Evaluation.

Result Command
GLUE bash scripts/eval_nlu.sh
MTEB v2 (SimCSE fine-tuning, then 41 tasks) bash scripts/eval_mteb.sh
BEIR (15 datasets) and MultiLongDocRetrieval bash scripts/eval_dpr.sh
CodeSearchNet and StackOverflowQA bash scripts/eval_code.sh
NanoBEIR bash scripts/eval_nanobeir.sh
Throughput vs. input length python src/nexterabert/model_speedbench.py
Masked-token loss vs. input length python src/nexterabert/model_pplbench.py

Other encoders go through the same stages, so a baseline can be evaluated under the identical protocol:

MODEL=answerdotai/ModernBERT-base bash scripts/eval_nlu.sh

6 Conclusion

With about one fifteenth of the pretraining tokens, NexteraBERT matches ModernBERT-base on GLUE and exceeds it on MTEB v2. It is several times faster on long inputs, and its maskedtoken loss stays stable up to eight times its longest training length. Even training with just 130B can achieve downstream task performance equal to or better than ModernBERT, so I think even small labs can conduct sufficient follow-up research.

Acknowledgements

My interest in this topic originated from reading the paper of ModernBERT, so I especially want to thank the authors, Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Nathan Cooper, Griffin Adams, Jeremy Howard, Iacopo Poli .

I thank the developers of transformers, python and pytorch.

I thank all the researchers for their efforts to date.

I thank Japan's high standard of education.

And most of all, thank you for your interest in this blog.

About us

Japanese independent researcher having shy and pampered personality. Twin-tail hair is a charm point. Interested in nlp. Usually using python and C.

Please contact us if you have any requests for joint research, writing, speaking engagements, or employment.

RikkaBotan_Logo

Reference

Community

Sign up or log in to comment