Sage Logo Sage

Offline-First Academic Intelligence.
Unlimited. 100% Private.

Sage bundles llama.cpp inference, ONNX embeddings, and open-weight LLMs into a standalone distribution. Choose a tier that matches your hardware and download.

Choose Your Distribution

Full tiers include pre-packaged models ready to run. Lite tiers ship the engine only, bring your own GGUF models.

x64 CPU

Sage Fast

CPU · Full Pack

3.42 GB

Standard CPU distribution with x64-optimized engine, Qwen-3.5 2B primary and 0.8B utility models pre-loaded.

Qwen-3.5 2B Primary
Qwen-3.5 0.8B Utility
Curriculum VectorDB
Engine llama.cpp CPU
Speed ~8-12 t/s
RAM 8 GB+
Download ZIP
CUDA 12.4

Sage Pro-Lite

GPU · Engine Only

2.14 GB

Minimal GPU runner with CUDA llama.cpp and ONNX pipelines. No models included.

No models included
Curriculum VectorDB
Engine llama.cpp CUDA
Speed ~35+ t/s
RAM 8 GB+
Download ZIP
x64 CPU

Sage Fast-Lite

CPU · Engine Only

1.61 GB

Minimal CPU runner with x64 llama.cpp and ONNX pipelines. No models included.

No models included
Curriculum VectorDB
Engine llama.cpp CPU
Speed ~8-12 t/s
RAM 8 GB+
Download ZIP

Tier Comparison

A quick look at what differs between distribution tiers.

Recommended
Pro
Fast
Pro-Lite
Fast-Lite
Backend CUDA 12.4 x64 CPU CUDA 12.4 x64 CPU
Primary Model Qwen-3.5 4B Qwen-3.5 2B
Utility Model Qwen-3.5 0.8B Qwen-3.5 0.8B
Download Size ~5.41 GB ~3.42 GB ~2.14 GB ~1.61 GB
Generation Speed ~35+ t/s ~8-12 t/s ~35+ t/s ~8-12 t/s
Min RAM 16 GB 8 GB 8 GB 8 GB
GPU Required NVIDIA CUDA No NVIDIA CUDA No