Jeffy v0.1.0-alpha.7 ships 13 CPU logistic regression classifiers over bge-large-en-v1.5 embeddings
Jeffy demonstrates practical linear probing on bge embeddings for thirteen classification tasks with documented accuracies. The approach prioritizes CPU execution and small deployable artifacts over peak accuracy on complex inference problems.
The package exposes an HTTP server and Python SDK that load a shared 1.2 GB bge-large-en-v1.5 encoder once, then apply task-specific linear heads. Thirteen classifiers ship with documented training splits; banking77 reaches 94.3 percent accuracy, snli 65.6 percent. Weights are stored as coefficients only, satisfying license terms listed in ATTRIBUTION.md. Custom training accepts CSV or JSONL and writes new heads to disk in seconds. Benchmark data in data/eval_results/benchmark.json shows F1 scores above 90 percent on four of the thirteen tasks and sub-70 percent on three. These numbers reflect held-out splits from the original datasets rather than external leaderboards. The gap versus fine-tuned transformers on SNLI and tweet sentiment matches expected results for linear probing of fixed embeddings. Operationally the design eliminates GPU requirements and model-size bloat for intent, spam, and topic routing workloads. Inference latency stays low because only a single embedding pass plus matrix multiply occur. Teams can version heads separately from the encoder, enabling on-device or air-gapped deployments where full transformer fine-tuning remains impractical.
Jeffy maintainers: repository reaches 300 GitHub stars and two external task submissions by March 2025
Sources (3)
- [1]Primary Source(https://github.com/nicobrenner/jeffy)
- [2]BGE Embeddings Paper(https://arxiv.org/abs/2307.09288)
- [3]Benchmark Results(https://github.com/nicobrenner/jeffy/blob/main/data/eval_results/benchmark.json)