TinyModels animated header



๐Ÿง  Tiny model. Tiny dataset. Real classification.

A compact few-shot spam classifier built with SetFit + BAAI/bge-small-en-v1.5.


โšก TINY โ†’ FAST โ†’ USEFUL

This model is a small experiment in few-shot text classification.

It learns to separate:

โœ‰๏ธ HAM โ€” legitimate email ๐Ÿšจ SPAM โ€” unwanted / suspicious email

The interesting part?

Only 16 labeled training examples.

                 โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                 โ”‚      16 EXAMPLES     โ”‚
                 โ”‚                      โ”‚
                 โ”‚   8 HAM  +  8 SPAM   โ”‚
                 โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                            โ”‚
                            โ–ผ
                โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                โ”‚ BAAI/bge-small-en-v1.5โ”‚
                โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                            โ”‚
                            โ–ผ
                   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                   โ”‚     SetFit     โ”‚
                   โ”‚  Few-shot NLP  โ”‚
                   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                           โ”‚
                           โ–ผ
                โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                โ”‚ Logistic Regression   โ”‚
                โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                            โ”‚
                     โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”
                     โ–ผ             โ–ผ
                   HAM           SPAM

๐Ÿ“ก Model Status

โš™๏ธ Component ๐Ÿ”ง Configuration
Task Spam Classification
Classes 2
Backbone BAAI/bge-small-en-v1.5
Framework SetFit
Dataset SetFit/enron_spam
Training Examples 16
Examples / Class 8
Iterations 20
Epochs 1
Batch Size 16
Classification Head Logistic Regression
Language English

๐Ÿ“Š Results

91.4% Accuracy

91.4% Macro F1


Accuracy   โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘  91.4%
Macro F1   โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘  91.4%

These results come from a very small few-shot training setup. They should not be interpreted as a benchmark against production spam-filtering systems.


๐Ÿงฌ The TinyModels Recipe

              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚   SetFit/enron    โ”‚
              โ”‚      _spam        โ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                        โ”‚
                  16 examples
                        โ”‚
                        โ–ผ
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚     BGE-small     โ”‚
              โ”‚   text encoder    โ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                        โ”‚
                   embeddings
                        โ”‚
                        โ–ผ
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚      SetFit       โ”‚
              โ”‚ contrastive loss  โ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                        โ”‚
                        โ–ผ
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚ LogisticRegressionโ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                        โ”‚
                 โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”
                 โ–ผ             โ–ผ
              โœ‰๏ธ HAM        ๐Ÿšจ SPAM

Training configuration

backbone: BAAI/bge-small-en-v1.5
dataset: SetFit/enron_spam

examples:
  total: 16
  per_class: 8

setfit:
  iterations: 20
  epochs: 1
  batch_size: 16
  loss: CosineSimilarityLoss
  distance_metric: cosine_distance

classifier:
  type: LogisticRegression

seed: 42

๐Ÿš€ Run It

pip install -q setfit
from setfit import SetFitModel

model = SetFitModel.from_pretrained(
    "TinyModels/setfit-banking-spam"
)

text = """
Congratulations! You have won $1,000,000.
Click here immediately to claim your prize.
"""

prediction = model(text)

print(prediction)

Example output

1

๐Ÿงช Quick Examples

๐Ÿšจ Spam

CONGRATULATIONS!!!

You have been selected to receive
$1,000,000. Click the link below
to claim your prize immediately.
โ†’ SPAM

โœ‰๏ธ Legitimate

Please find attached the global markets
monitor for the week ending 12 January 2001.
โ†’ HAM

๐Ÿง  Why SetFit?

Traditional supervised classification can require a large labeled dataset.

SetFit takes a different route:

        Large Dataset
             โœ•
             โ”‚
             โ”‚
       โ”Œโ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”
       โ”‚  SetFit   โ”‚
       โ””โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”˜
             โ”‚
             โ–ผ
     Few labeled examples
             โ”‚
             โ–ผ
       Useful classifier

This makes the experiment useful for exploring:

  • โšก Few-shot learning
  • ๐Ÿง  Sentence embeddings
  • ๐Ÿ“š Text classification
  • ๐Ÿšจ Spam detection
  • ๐Ÿ”ฌ Efficient training
  • ๐Ÿค— SetFit

๐Ÿงฉ Architecture

INPUT
  โ”‚
  โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚  BAAI/bge-small-en-v1.5     โ”‚
โ”‚                             โ”‚
โ”‚  Sentence Transformer       โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
               โ”‚
               โ–ผ
        Dense Embedding
               โ”‚
               โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚    Logistic Regression      โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
               โ”‚
        โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”
        โ–ผ             โ–ผ
      HAM           SPAM
       โœ‰๏ธ             ๐Ÿšจ

โš ๏ธ Limitations

This is intentionally a tiny experimental model.

Because only 16 examples were used for training:

  • Performance can vary on unseen data.
  • Domain shift can significantly affect predictions.
  • Unusual spam may be missed.
  • Enron-style email does not represent every modern spam pattern.
  • The reported score comes from a lightweight few-shot experiment.

Do not use this model as the sole component of a security-critical email filtering system.


๐Ÿ“ฆ Model Identity

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚              TINYMODEL               โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚                                      โ”‚
โ”‚  MODEL      SetFit Banking Spam      โ”‚
โ”‚  BACKBONE   BGE-small                โ”‚
โ”‚  TASK       Binary Classification    โ”‚
โ”‚  DATA       Enron Spam               โ”‚
โ”‚  EXAMPLES   16                       โ”‚
โ”‚  RESULT     91.4% Accuracy           โ”‚
โ”‚                                      โ”‚
โ”‚  STATUS     โ— EXPERIMENTAL           โ”‚
โ”‚                                      โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

โšก 16 Examples.

๐Ÿง  One Small Model.

๐Ÿšจ One Real Task.




TinyModels โ€” building small models that actually do things.

Downloads last month
11
Safetensors
Model size
33.4M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for TinyModels/Setfit-Banking-Spam

Finetuned
(521)
this model

Dataset used to train TinyModels/Setfit-Banking-Spam