SentenceTransformer based on sentence-transformers/all-MiniLM-L6-v2

This is a sentence-transformers model finetuned from sentence-transformers/all-MiniLM-L6-v2. It maps sentences & paragraphs to a 384-dimensional dense vector space and can be used for retrieval.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Base model: sentence-transformers/all-MiniLM-L6-v2
  • Maximum Sequence Length: 256 tokens
  • Output Dimensionality: 384 dimensions
  • Similarity Function: Cosine Similarity
  • Supported Modality: Text

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'architecture': 'BertModel'})
  (1): Pooling({'embedding_dimension': 384, 'pooling_mode': 'mean', 'include_prompt': True})
  (2): Normalize({})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
sentences = [
    'What is 100 degrees Fahrenheit converted to Celsius?',
    'Architect a globally distributed microservices deployment platform enforcing zero-trust networking, canary releases with automated rollback based on multi-region SLOs, cross-cloud secrets rotation, and infrastructure cost allocation across AWS, Azure, and GCP.',
    'What is the default network port for HTTP?',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 384]

# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[ 1.0000, -0.2282,  0.9992],
#         [-0.2282,  1.0000, -0.2242],
#         [ 0.9992, -0.2242,  1.0000]])

Training Details

Training Dataset

Unnamed Dataset

  • Size: 871 training samples
  • Columns: sentence and label
  • Approximate statistics based on the first 871 samples:
    sentence label
    type string int
    details
    • min: 6 tokens
    • mean: 34.77 tokens
    • max: 179 tokens
    • 0: ~19.86%
    • 1: ~28.47%
    • 2: ~27.10%
    • 3: ~24.57%
  • Samples:
    sentence label
    Write a discharge instruction template for patients recovering from total knee replacement surgery, including medication guidelines, physical therapy milestones, and red-flag symptoms. 1
    Prove that alpha-beta pruning in game tree search returns identical minimax values as full search for deterministic two-player zero-sum games with perfect information. 3
    Explain the zone of proximal development using a specific classroom learning scenario. 1
  • Loss: BatchAllTripletLoss with these parameters:
    {
        "distance_metric": "euclidean_distance",
        "margin": 5
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 32
  • num_train_epochs: 20
  • learning_rate: 2e-05
  • warmup_steps: 0.1
  • weight_decay: 0.01
  • batch_sampler: group_by_label

All Hyperparameters

Click to expand
  • per_device_train_batch_size: 32
  • num_train_epochs: 20
  • max_steps: -1
  • learning_rate: 2e-05
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_steps: 0.1
  • optim: adamw_torch_fused
  • optim_args: None
  • weight_decay: 0.01
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • optim_target_modules: None
  • gradient_accumulation_steps: 1
  • average_tokens_across_devices: True
  • max_grad_norm: 1.0
  • label_smoothing_factor: 0.0
  • bf16: False
  • fp16: False
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • use_liger_kernel: False
  • liger_kernel_config: None
  • use_cache: False
  • neftune_noise_alpha: None
  • torch_empty_cache_steps: None
  • auto_find_batch_size: False
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • include_num_input_tokens_seen: no
  • log_level: passive
  • log_level_replica: warning
  • disable_tqdm: False
  • project: huggingface
  • trackio_space_id: None
  • trackio_bucket_id: None
  • trackio_static_space_id: None
  • per_device_eval_batch_size: 8
  • prediction_loss_only: True
  • eval_on_start: False
  • eval_do_concat_batches: True
  • eval_use_gather_object: False
  • eval_accumulation_steps: None
  • include_for_metrics: []
  • batch_eval_metrics: False
  • save_only_model: False
  • save_on_each_node: False
  • enable_jit_checkpoint: False
  • push_to_hub: False
  • hub_private_repo: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_always_push: False
  • hub_revision: None
  • load_best_model_at_end: False
  • ignore_data_skip: False
  • restore_callback_states_from_checkpoint: False
  • full_determinism: False
  • seed: 42
  • data_seed: None
  • use_cpu: False
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • dataloader_prefetch_factor: None
  • remove_unused_columns: True
  • label_names: None
  • train_sampling_strategy: random
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • ddp_static_graph: None
  • ddp_backend: None
  • ddp_timeout: 1800
  • fsdp: []
  • fsdp_config: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}
  • deepspeed: None
  • debug: []
  • skip_memory_metrics: True
  • do_predict: False
  • resume_from_checkpoint: None
  • warmup_ratio: None
  • local_rank: -1
  • prompts: None
  • batch_sampler: group_by_label
  • multi_dataset_batch_sampler: proportional
  • router_mapping: {}
  • learning_rate_mapping: {}

Training Logs

Epoch Step Training Loss
0.3704 10 4.9309
0.7407 20 4.9160
1.1111 30 4.8961
1.4815 40 4.7981
1.8519 50 4.6499
2.2222 60 4.4857
2.5926 70 4.2761
2.9630 80 4.1235
3.3333 90 4.0387
3.7037 100 3.9355
4.0741 110 3.9115
4.4444 120 3.8240
4.8148 130 3.7550
5.1852 140 3.7214
5.5556 150 3.6528
5.9259 160 3.6749
6.2963 170 3.6371
6.6667 180 3.6124
7.0370 190 3.5848
7.4074 200 3.5935
7.7778 210 3.5504
8.1481 220 3.5539
8.5185 230 3.5427
8.8889 240 3.5122
9.2593 250 3.4894
9.6296 260 3.5147
10.0 270 3.4909
10.3704 280 3.5146
10.7407 290 3.4988
11.1111 300 3.4591
11.4815 310 3.4997
11.8519 320 3.4847
12.2222 330 3.4704
12.5926 340 3.4826
12.9630 350 3.4474
13.3333 360 3.4807
13.7037 370 3.4772
14.0741 380 3.4373
14.4444 390 3.4747
14.8148 400 3.4712
15.1852 410 3.4310
15.5556 420 3.4720
15.9259 430 3.4650
16.2963 440 3.4394
16.6667 450 3.4700
17.0370 460 3.4330
17.4074 470 3.4684
17.7778 480 3.4706
18.1481 490 3.4456
18.5185 500 3.4681
18.8889 510 3.4560
19.2593 520 3.4415
19.6296 530 3.4683
20.0 540 3.4167

Training Time

  • Training: 12.6 minutes

Framework Versions

  • Python: 3.11.13
  • Sentence Transformers: 5.4.1
  • Transformers: 5.8.0
  • PyTorch: 2.11.0+cpu
  • Accelerate: 1.13.0
  • Datasets: 4.8.5
  • Tokenizers: 0.22.2

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

BatchAllTripletLoss

@misc{hermans2017defense,
    title={In Defense of the Triplet Loss for Person Re-Identification},
    author={Alexander Hermans and Lucas Beyer and Bastian Leibe},
    year={2017},
    eprint={1703.07737},
    archivePrefix={arXiv},
    primaryClass={cs.CV}
}
Downloads last month
58
Safetensors
Model size
22.7M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cnuland/semantic-claw-router-embeddings-v1

Papers for cnuland/semantic-claw-router-embeddings-v1