# Training Sentence Transformers with Softmax Loss

**URL:** <https://community.pinecone.io/t/training-sentence-transformers-with-softmax-loss/88>\
**Category:** General\
**Created:** [February 1, 2022, 7:59pm UTC](https://community.pinecone.io/t/training-sentence-transformers-with-softmax-loss/88 "2022-02-01T19:59:49Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![discobot](https://sea2.discourse-cdn.com/flex020/user_avatar/community.pinecone.io/discobot/32/2_2.png) [@discobot](https://community.pinecone.io/u/discobot)\
**Post date:** [February 1, 2022, 7:59pm UTC](https://community.pinecone.io/t/training-sentence-transformers-with-softmax-loss/88/1 "2022-02-01T19:59:49Z")

</div>

Our article introducing [sentence embeddings and transformers](https://www.pinecone.io/learn/sentence-embeddings/) explained that these models can be used across a range of applications, such as semantic textual similarity (STS), semantic clustering, or information retrieval (IR) using concepts rather than words.

This article dives deeper into the training process of the first sentence transformer, _sentence-BERT_, or more commonly known as _SBERT_. We will explore the **N** atural **L** anguage **I** nference (NLI) training approach of _softmax loss_ to fine-tune models for producing sentence embeddings.

* * *
This is a companion discussion topic for the original entry at [https://www.pinecone.io/learn/train-sentence-transformers-softmax/](https://www.pinecone.io/learn/train-sentence-transformers-softmax/)
