Offline-First Academic Intelligence.
Unlimited.
100% Private.
Sage bundles llama.cpp inference, ONNX embeddings, and open-weight LLMs into a standalone distribution. Choose a tier that matches your hardware and download.
Choose Your Distribution
Full tiers include pre-packaged models ready to run. Lite tiers ship the engine only, bring your own GGUF models.
Sage Pro
GPU · Full Pack
Flagship GPU distribution with CUDA 12.4 backend, Qwen-3.5 4B primary and 0.8B utility models pre-loaded.
Sage Fast
CPU · Full Pack
Standard CPU distribution with x64-optimized engine, Qwen-3.5 2B primary and 0.8B utility models pre-loaded.
Sage Pro-Lite
GPU · Engine Only
Minimal GPU runner with CUDA llama.cpp and ONNX pipelines. No models included.
Sage Fast-Lite
CPU · Engine Only
Minimal CPU runner with x64 llama.cpp and ONNX pipelines. No models included.
Tier Comparison
A quick look at what differs between distribution tiers.
|
Recommended
Pro
|
Fast
|
Pro-Lite
|
Fast-Lite
|
|
|---|---|---|---|---|
| Backend | CUDA 12.4 | x64 CPU | CUDA 12.4 | x64 CPU |
| Primary Model | Qwen-3.5 4B | Qwen-3.5 2B | — | — |
| Utility Model | Qwen-3.5 0.8B | Qwen-3.5 0.8B | — | — |
| Download Size | ~5.41 GB | ~3.42 GB | ~2.14 GB | ~1.61 GB |
| Generation Speed | ~35+ t/s | ~8-12 t/s | ~35+ t/s | ~8-12 t/s |
| Min RAM | 16 GB | 8 GB | 8 GB | 8 GB |
| GPU Required | NVIDIA CUDA | No | NVIDIA CUDA | No |