GGUF quantizations of Ling-3.0-tiny with 256K context, bilingual ES/EN calibration, and architecture-aware mixed precision

Ling-3.0-tiny is a mixture-of-experts model by inclusionAI (MIT license) with 7.9B total parameters and approximately 1.3B active per token. The architecture, bailingmoe3, is hybrid: 18 blocks use KDA (Kimi Delta Attention) with a fixed-size recurrent state, and 6 blocks (3, 7, 11, 15, 19, 23) use MLA (Multi-head Latent Attention) with a compressed KV latent. Each MoE layer has 128 routed experts, 8 active per token, plus one shared expert active on every token. Block 0 uses a dense FFN instead of MoE. This block layout matches the officially documented 3:1 alternating stacking of KDA and MLA layers.

The base model includes a native context of 131,072 tokens, within an architecture that claims to be designed to support up to 1M. The maintainers' recommended deployment recipe extends the model to 262,144 tokens via YaRN, with factor 2.0 over the native rope_theta of 6,000,000, and their Terminal-Bench 2.1 evaluation was conducted at that window, making 256K the validated operating point of the model. Community GGUF quants, however, ship at the native window only. This release embeds the same officially recommended YaRN configuration directly into the GGUF metadata, so the validated 262,144-token window is available to llama.cpp users without runtime flags. In addition, the importance matrix is calibrated on a bilingual Spanish/English dataset, and a per-tensor precision map is applied, grounded in how quantization error propagates through this specific hybrid architecture.

Quantized with llama.cpp 0.4.0-dev, build 10902, commit df03399b8.

Original model: https://huggingface.co/inclusionAI/Ling-3.0-tiny

Model details

Property Value
Parameter count 7.9B total, ~1.3B active per token
Architecture bailingmoe3, hybrid KDA + MLA + MoE
KDA blocks 18 (blocks 0–2, 4–6, 8–10, 12–14, 16–18, 20–22)
MLA blocks 6 (blocks 3, 7, 11, 15, 19, 23)
Experts per MoE layer 128 routed, 8 active, 1 shared
Context 262,144 tokens (extended from 131,072 via YaRN, architecture supports up to 1M)
Input text
imatrix yes, bilingual ES/EN, ~1.67M tokens
Perplexity measured yes, Spanish and English held-out sets

Prompt format

<role>SYSTEM</role>{system_prompt}
detailed thinking on<|role_end|><role>HUMAN</role>{prompt}<|role_end|><role>ASSISTANT</role>

The template is embedded in the files and applied automatically by llama-server --jinja.

Which file to choose. For most use cases Q4_K_M (4.89 GB) is the practical default; it measures within 0.3% of BF16 perplexity in Spanish. Q5_K_M (5.76 GB) is the quality sweet spot of this ladder: the only level that matches or improves on BF16 in both languages, at 36% of the original size, a non-monotonic result consistent with what Kurt (2026) documented for 5-bit formats on Llama-3.1-8B. IQ4_XS (4.67 GB) is smaller than Q4_K_M while measuring better in English.

Available files

All files embed the 256K YaRN extension, the bilingual importance matrix, and the per-tensor precision map.

Filename Quant type File size Approx. VRAM Description
Ling-3.0-tiny-256K-bf16.gguf bf16 15.80GB 16 GB or CPU Full BF16 weights with 256K YaRN metadata. Reference baseline.
Ling-3.0-tiny-256K-Q8_0.gguf Q8_0 7.9GB 10–12 GB Near-lossless. Measures at parity with Q6_K; generally unnecessary.
Ling-3.0-tiny-256K-Q6_K.gguf Q6_K 6.3GB 8–10 GB Very high quality. Measures statistically identical to Q8_0 in both languages.
Ling-3.0-tiny-256K-Q5_K_M.gguf Q5_K_M 5.4GB 8 GB Best quality/size ratio. Only level matching BF16 in both languages. Recommended.
Ling-3.0-tiny-256K-Q4_K_M.gguf Q4_K_M 4.6GB 6–8 GB Practical default. +0.3% ES / +1.7% EN vs BF16. Recommended.
Ling-3.0-tiny-256K-IQ4_XS.gguf IQ4_XS 4.4GB 6 GB 0.22 GB smaller than Q4_K_M with better EN perplexity. Recommended for tight VRAM.
Ling-3.0-tiny-256K-IQ3_M.gguf IQ3_M 3.7GB 4–6 GB Lowest level with controlled degradation. +3.9% ES / +1.6% EN vs BF16.
Ling-3.0-tiny-256K-IQ2_M.gguf IQ2_M 2.9GB 4 GB Not recommended for reasoning tasks. +20.6% ES / +9.3% EN vs BF16.

VRAM figures are for weights at moderate context. At the full 262,144-token window, budget approximately 1.8 GB additional for the KV cache in f16: only the six MLA layers carry a cache, since the KDA layers hold a fixed-size recurrent state that does not grow with context length. The cache can be reduced with --cache-type-k q8_0.

Downloading with the Hugging Face CLI

Click to expand
pip install -U "huggingface_hub[cli]"
hf download stornic56/Ling-3.0-tiny-256K-GGUF \
  --include "Ling-3.0-tiny-256K-Q4_K_M.gguf" \
  --local-dir ./

How to run

These quants run with llama.cpp, installable via llama.app:

curl -LsSf https://llama.app/install.sh | sh
llama-server -hf stornic56/Ling-3.0-tiny-256K-GGUF:Q4_K_M

To use the full 256K context, no --rope-scaling flags are needed. The YaRN parameters are embedded in the GGUF metadata; users only request the window size:

llama-server -m Ling-3.0-tiny-256K-Q4_K_M.gguf -c 262144 -fa --jinja

-fa (flash attention) is strongly recommended at long context. --jinja enables the native chat template including tool-calling.

The bailingmoe3 architecture requires llama.cpp b10470 or newer; these files were produced with b10902.

The model card recommends temperature=1.0, top_p=0.95, and top_k=20. Thinking mode is enabled by default by the chat template.

Compatible with: LM Studio, koboldcpp, ramalama, Jan AI, and Text Generation Web UI.

llama.cpp repacks weights at load time for faster inference on ARM and AVX machines, so no special quant choice is required for CPU inference.


Quantization methodology

This is not a standard single-pass imatrix quantization. Three things distinguish it from the community default: the context window is extended before conversion, the importance matrix is calibrated on a bilingual dataset, and a per-tensor precision map is applied based on how quantization error propagates through this specific architecture.

Context extension to 262,144 tokens with YaRN

The base model is trained with a 131,072-token context. Standard GGUF conversions carry no rope-scaling metadata, which limits every derived quant to the native window.

Extension was applied at conversion time by setting max_position_embeddings = 262144 and adding a rope_scaling block (rope_type: yarn, factor: 2.0, original_max_position_embeddings: 131072) to the model configuration, preserving the native rope_theta of 6,000,000. convert_hf_to_gguf.py then writes the following metadata, which every quant in this repo inherits from the BF16 source:

bailingmoe3.context_length = 262144
bailingmoe3.rope.scaling.type = yarn
bailingmoe3.rope.scaling.factor = 2.0
bailingmoe3.rope.scaling.original_context_length = 131072
bailingmoe3.rope.scaling.yarn_beta_fast = 32.0
bailingmoe3.rope.scaling.yarn_beta_slow = 1.0
bailingmoe3.rope.freq_base = 6000000.0

Two controls validated the extension. First, an attribution control: the unextended community reference run with runtime YaRN flags (--rope-scaling yarn --rope-scale 2.0) produces perplexity statistically identical to the metadata path (6.849 vs 6.856 in Spanish at Q8_0), confirming the embedded metadata is applied and equivalent to the runtime path. Second, a short-context control: at 4096-token evaluation the extended model measures approximately 2% lower perplexity in Spanish and 5% lower in English than the unextended model, so the extension does not penalize short-context use at this scale factor.

Method: Peng, Quesnelle, Fan, and Shippole (2023), YaRN: Efficient Context Window Extension of Large Language Models, arXiv:2309.00071.

A third point of validation comes from the model card itself. The maintainers' recommended SGLang deployment recipe applies the identical rope_scaling configuration at runtime, with rope_type yarn, factor 2.0, rope_theta 6000000, partial_rotary_factor 0.5, and original_max_position_embeddings 131072, at a context length of 262,144, and their Terminal-Bench 2.1 evaluation protocol runs at that window. The metadata embedded in these files reproduces that official configuration, making it the default for llama.cpp users rather than a per-deployment runtime setting.

The injected metadata can be verified directly:

python3 -m gguf.scripts.gguf_dump Ling-3.0-tiny-256K-Q4_K_M.gguf | grep -iE "rope|context"

Importance matrix calibration on a bilingual dataset

The importance matrix was computed with llama-imatrix from build 10902 over the BF16 256K source, covering 3,282 chunks and producing 332 importance matrix entries with all tensors covered. The imatrix file is published alongside the quants.

llama-imatrix -m Ling-3.0-tiny-256K-BF16.gguf \
  -f calibracion_final.txt \
  -o imatrix_main.gguf \
  --parse-special --no-ppl

The calibration dataset contains approximately 1.67M tokens (1,675,111 measured with the model tokenizer) in a Spanish/English mixture:

Block Share Sources
Tool-calling, EN and ES ~36% interstellarninja/hermes_reasoning_tool_use, Salesforce/xlam-function-calling-60k, bartowski calibration v6 conversations
Spanish conversational/instruction ~22% OpenAssistant OASST1 (ES), Dolly (ES), Aya (ES), BSC MentorES
Spanish academic/medical ~15% Multilingual BioASQ, MedExpQA, Spanish Wikipedia
English prose ~17% bartowski calibration v6 prose, eaddario/imatrix-calibration
Code and mathematics ~10% GitHub code, MetaMathQA

The bilingual composition is intentional. The model is aimed at deployment in Spanish and English, and English-only calibration is a suboptimal default for that use. Two findings support the choice. Chimoto, Elhoushi, and Bassett (2026) show that non-English and mixed-language calibration sets consistently outperform English-only baselines for activation-based quantizers, with balanced mixtures the most robust across languages. Liu, Sun, Zhang, et al. (2025) independently show that the composition of calibration data has a measurable effect on quantized model quality, with domain-matched calibration improving quantized reasoning accuracy by up to 9.8 percentage points in their experiments. The effect was measured directly on this model: at Q4_K_M under uniform precision assignment, the bilingual imatrix yields 3.1% lower Spanish perplexity than the English-only reference imatrix applied to the same BF16 base. This benefit becomes fully visible only after the per-tensor precision map is applied; see the measured results for the complete picture.

Per-tensor precision map

The architecture is not a plain Transformer, and standard quantization recipes were designed for dense Transformer models. These quants apply a precision map based on how error propagates in this specific hybrid. Three classes of quantization error motivate the assignments.

Type 1, isolated error. The error affects the current forward pass only; it is not stored or re-read across the sequence. MoE redundancy and layer depth absorb the noise. This describes routed experts (93.75% dormant per token by design), embeddings, and output projections.

Type 2, accumulated error. The error is written into a state that is re-read at every subsequent token. At 262,144 tokens, a bias introduced at token 100 remains in the representation at token 200,000. This describes the KV latent path (attn_kv_a_mqa, attn_k_b, attn_v_b, the tensors the KV cache stores and retrieves) and the KDA write path (attn_k, attn_v, which feed the recurrent delta-rule state). Zhao, Shen, Shi, et al. (2025) document this pattern in their analysis of DeepSeek MoE quantization, establishing that components on the always-active path require higher precision than routed experts.

Type 3, categorical error. The error does not add proportional noise; it pushes a sigmoid or softmax across its decision boundary, changing which expert is selected or how much state is retained. A small quantization error in a gate or router weight can corrupt expert selection for every token. The GLM-5.3-Flash GGUF release by the Qtum Foundation documents this for the router: "MoE router; compressing it routes to the wrong experts." For the recurrent state specifically, Zhang, Tan, Sun, et al. (2026) present the first study of post-training quantization of recurrent states in GDN and KDA based language models, establishing it as a distinct problem: these states are conventionally kept in FP32, and their quantization requires dedicated decay-aware treatment. F32 assignment on the state controls follows that convention.

A fourth practical consideration is tensor shape. ssm_beta.weight and attn_gate.weight have only 16 output dimensions. Quantization logs confirm they convert to Q8_0 (and to IQ3_S in IQ runs) without explicit rules, because a row width of 16 is incompatible with the 256-weight super-blocks of K-quants and the lattice codebooks of IQ-quants. With the explicit F32 override, the fallback warnings disappear from every IQ run.

One assignment is tapered by level rather than uniform. The KDA attention projections are held at Q8_0 down to IQ4_XS, but drop to Q6_K at IQ3_M and IQ2_M. At those targets the experts are already at roughly 3.7 and 2.7 bits per weight respectively, and carrying 54 tensors at Q8_0 would push the file size disproportionately, an estimated ~0.13 GiB per level. The measured results indicate the compromise holds for IQ3_M (+3.9% ES, controlled) but IQ2_M crosses the quality cliff regardless, with the taper contributing to but not solely responsible for that.

The resulting map:

Component Assignment Rationale
KDA recurrent-state controls (ssm_a, ssm_beta, ssm_dt, ssm_conv1d_*, ssm_norm) F32 Type 3. Controls decay and write-strength of the recurrent state; recurrent-state quantization in GDN/KDA is a distinct, only recently studied problem (Zhang et al., 2026). The default path leaves ssm_beta at Q8_0, verified in quantization logs; this map corrects that.
MoE router (ffn_gate_inp, exp_probs_b) F32 Type 3. Routing errors corrupt expert selection for every token. Already F32-native in llama-quantize for most formats; anchored explicitly for robustness.
MLA gate (attn_gate) F32 Type 3. Selective gate at the 6 MLA blocks. Same class as ssm_beta: 16-dim weight, incompatible with standard quantization blocks, ~0.6 MB total cost.
All normalization weights (*_norm) F32 Type 3. RMSNorm parameter spikes propagate through the entire model. Already F32-native; anchored for completeness.
MLA KV latent path (attn_kv_a_mqa, attn_k_b, attn_v_b) Q8_0 Type 2. This is what the KV cache stores and retrieves; degrading the latent compression degrades every retrieved K/V over the full context. ~34 MB total.
KDA attention projections (attn_k, attn_v, attn_q) Q8_0 at Q8/Q6/Q5/Q4 and IQ4_XS; Q6_K at IQ3_M and IQ2_M Type 2 for k/v: they feed the delta-rule recurrent state, and noise introduced before the state write persists in the state. Type 1 for q: recomputed per token from the current token only. The taper at the two most compressed IQ levels trades part of this protection for size; see the note above.
MLA query projections (attn_q_a, attn_q_b) Q8_0 at Q8/Q6; Q6_K at Q5/Q4 Type 1. Recomputed at each step, not cached. Tapered at lower levels.
Shared expert (ffn_*_shexp) Q8_0 Active on every token without exception, unlike the 8/128 routed experts. Zhao et al. (2025) establish that shared components require higher precision than routed experts, and the Qtum Foundation map for GLM-5.3-Flash protects the shared expert identically.
Dense FFN layer (blk.0.ffn_*) Q8_0 The single dense MLP layer, active on every token; same logic as the shared expert. The Qtum Foundation map applies the same treatment to the dense MLP layers of GLM-5.3-Flash.
Attention output projections (attn_output) Q8_0 at Q8/Q6; Q6_K at Q5/Q4 Type 1. Output projection, not part of the recurrent write path.
SSM input gates (ssm_f_a, ssm_g_a) Q6_K Input projections into the gate mechanism, distinct from the state controls themselves (which stay F32). They feed the recurrence but are not part of the stored state.
Embeddings and output head (token_embd, output.weight) Q8_0 Low redundancy: a vocabulary of 157,184 tokens with embedding dimension 1,536 leaves little margin to absorb quantization noise per row. Each row is unique; the error cannot average out.
Routed experts (ffn_*_exps) base level of each quant Type 1, diluted. 93.75% of experts are dormant per token by design; error across the 8 active experts averages out rather than composing. Zhao et al. (2025) demonstrate that routed experts in DeepSeek-scale MoE tolerate the most aggressive compression of any component.

A representative quantization command:

llama-quantize \
  --imatrix imatrix_main.gguf \
  --tensor-type-file tensor-types-Q4_K_M.txt \
  --output-tensor-type q8_0 \
  --token-embedding-type q8_0 \
  Ling-3.0-tiny-256K-BF16.gguf \
  Ling-3.0-tiny-256K-Q4_K_M.gguf \
  Q4_K_M

The complete tensor type maps for all seven levels are published in the reproducibility/ directory of this repo, together with the validation script. A validation script (validate_tensor_types.py) is also included: it reads the GGUF with gguf.GGUFReader, verifies the KDA/MLA/FFN block map against the expected pattern, checks for conflicting rules, and reports any patterns that match zero tensors. This is how a lookahead-regex bug in an early version of the maps was caught before any compute was spent.

A parallel application of this approach exists for a larger model of the same architectural family. The GLM-5.3-Flash GGUF release by the Qtum Foundation applies the same pattern to a 320 billion parameter hybrid model with 34 KDA layers, 11 sparse attention layers, 288 routed experts, and one shared expert: router in F32, shared expert in Q8_0, KDA recurrent-state controls in F32, and dense MLP layers in Q8_0. The convergence of both efforts on identical protection patterns for the same architectural components, arrived at independently, supports the robustness of the assignment.

The empirical validation of this map on this model is the compression of Q4_K_M degradation from +8.6% under the standard single-precision pipeline to +0.3% in Spanish, and from +5.7% to +1.7% in English, measured against the BF16 baseline on held-out test sets.

Measured results

Perplexity was measured with llama-perplexity on held-out test sets (~57K tokens per language, 14 windows of 4,096 tokens, -c 4096 --chunks 30). The Spanish set is general-domain; the English set mixes general and technical content. Absolute values are not comparable across languages; within-language comparisons under the same evaluation are valid. Differences below 1% fall within measurement variance at this window count.

Quant PPL ES ± PPL EN ± ES vs BF16 EN vs BF16
BF16 7.786 0.128 20.909 0.404 reference reference
Q8_0 7.728 0.127 20.798 0.401 −0.7% −0.5%
Q6_K 7.723 0.127 20.795 0.401 −0.8% −0.5%
Q5_K_M 7.790 0.128 20.660 0.398 +0.1% −1.2%
Q4_K_M 7.813 0.128 21.256 0.411 +0.3% +1.7%
IQ4_XS 7.958 0.131 20.845 0.400 +2.2% −0.3%
IQ3_M 8.090 0.132 21.245 0.407 +3.9% +1.6%
IQ2_M 9.393 0.155 22.853 0.436 +20.6% +9.3%

Q5_K_M measures at or below BF16 in both languages. The apparent improvement is within the margin of error at 14 windows and should be read as statistical parity, not a real gain over the source. Kurt (2026) documents the same phenomenon on Llama-3.1-8B-Instruct, where the 5-bit formats Q5_0 and Q5_1 slightly exceed the F16 baseline on the aggregate downstream mean, and explicitly cautions that "modest improvements can reflect noise or idiosyncrasies of the scoring pipeline rather than true general superiority." The practical takeaway is that Q5_K_M achieves damage indistinguishable from zero at 36% of the original size.

IQ4_XS vs Q4_K_M. IQ4_XS (4.25 bpw effective on routed experts, guided by the importance matrix with codebook encoding) measures better than Q4_K_M (4.5 bpw K-quant on the same experts) in English while being 0.22 GB smaller, at the cost of slightly worse Spanish perplexity. The importance matrix provides a larger benefit to IQ formats than to K-quants, where it only influences scale selection, consistent with what Sparrenberg, Deußer, Berger, and Sifa (2025) document for IQ-quants in llama.cpp.

External reference (same test sets, same evaluation parameters): the leading community quantization of this model (bartowski, Q4_K_M, native 131K context, English-only calibration) measures 8.456 in Spanish and 22.093 in English. The Q4_K_M in this release measures 7.813 and 21.256, which is 7.6% and 3.8% lower respectively, at a slightly smaller file size (4.89 vs 4.92 GB). The comparison is not fully controlled: this release uses a 256K-extended base with YaRN, a different importance matrix, and the per-tensor precision map. All three factors contribute to the difference; they cannot be separated without additional ablations.

Honest notes and limitations

English technical content. Imatrix-based quantization of this model measures approximately 2% higher perplexity than uncalibrated quantization on raw arXiv-style text. This was verified to be a structural property of importance matrices on sparse-expert architectures: two importance matrices of opposite corpus composition (the bilingual one used here and the English-only reference, both applied to the same BF16 base) produced essentially the same effect (22.93 vs 22.94 EN perplexity). The per-tensor map recovers most of the gap. No functional difference between imatrix and no-imatrix was found across deterministic evaluations of mathematics, code, and tool-calling via llama-server --jinja with real schemas.

Q8_0 and Q6_K are Pareto-dominated by Q5_K_M in these measurements. Carrying a larger file for no measured quality gain is a valid choice when maximum precision is the goal, but Q5_K_M is the practical quality ceiling of this ladder.

IQ2_M is below the quality threshold for reasoning tasks. It is included for completeness; the model still produces coherent output, but the +20.6% Spanish perplexity and the literature on quantized reasoning models (Liu et al., 2025, who find that harder tasks suffer up to 4× greater degradation and that 3-bit weight-only quantization becomes risky while 4-bit remains nearly lossless) both indicate that compression at this level degrades multi-step reasoning disproportionately.

Perplexity is not a complete proxy for downstream quality. Kurt (2026) documents cases where quantization formats with identical perplexity diverge meaningfully on instruction-following and reasoning benchmarks, and concludes that intrinsic metrics alone are insufficient to select deployment defaults. The test sets here are held-out general-domain text; performance on domain-specific tasks was not benchmarked beyond the deterministic functional evaluation described above.


Supporting literature

The precision assignments and calibration approach in this release are grounded in the following work.

Quantization methods and behavior in llama.cpp:

  • Sparrenberg, Deußer, Berger, and Sifa (2025). Small and Fast LLMs on Commodity Hardware: Post-Training Quantization in llama.cpp. IEEE DSAA. DOI: 10.1109/DSAA65442.2025.11247985. Survey of PTQ methods in llama.cpp including K-quants and IQ-quants; documents that IQ-quants benefit more from a good importance matrix than K-quants.
  • Kurt (2026). Which Quantization Should I Use? A Unified Evaluation of llama.cpp Quantization on Llama-3.1-8B-Instruct. arXiv:2601.14277. Unified evaluation of 13 GGUF formats; documents the 5-bit at-or-above-FP16 phenomenon, GSM8K sensitivity to 3-bit compression (−9.32 points at Q3_K_S), and that quantization format matters beyond nominal bit-width.

Quantization of reasoning models:

  • Liu, Sun, Zhang, Bai, Yu, Yu, Yuan, and Hou (2025). Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models. COLM 2025. arXiv:2504.04823. First systematic study of quantized reasoning models; finds that harder tasks suffer up to 4× greater degradation, that 4-bit weight-only quantization is nearly lossless while 3-bit becomes risky, and that calibration data composition has a measurable effect on quantized quality.

MoE quantization behavior:

  • Zhao, Shen, Shi, Huang, Chen, and Wang (2025). Quantitative Analysis of Performance Drop in DeepSeek Model Quantization. arXiv:2505.02390. Establishes that routed experts tolerate aggressive compression while shared components and attention require higher precision. Basis for the experts-at-base-level, shared-expert-at-Q8_0 strategy.
  • Frantar, Ashkboos, Hoefler, and Alistarh (2023). GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers. ICLR 2023. arXiv:2210.17323. Theoretical basis for the importance matrix as a diagonal approximation of Hessian curvature per row.

Recurrent-state quantization:

  • Zhang, Tan, Sun, Yu, Jiang, Xie, Cai, and Zeng (2026). DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization. arXiv:2608.27513. First study of post-training quantization of recurrent states in GDN and KDA based language models; establishes recurrent-state quantization as a distinct problem and validates the convention of keeping state controls in FP32.

Bilingual calibration:

  • Chimoto, Elhoushi, and Bassett (2026). Calibrating Beyond English: Language Diversity for Better Quantized Multilingual LLMs. arXiv:2601.18306. Shows that non-English and mixed-language calibration sets consistently outperform English-only baselines for activation-based quantizers.

Context extension:

  • Peng, Quesnelle, Fan, and Shippole (2023). YaRN: Efficient Context Window Extension of Large Language Models. ICLR 2024. arXiv:2309.00071. Method used for the context extension.

Credits

Downloads last month
930
GGUF
Model size
8B params
Architecture
bailingmoe3
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for stornic56/Ling-3.0-tiny-256K-GGUF

Quantized
(31)
this model

Papers for stornic56/Ling-3.0-tiny-256K-GGUF