Offline AI · On-Device

Local LLM, running entirely offline.
Switch to Claude, GPT or Grok in the same chat.

AiMixup is the only workspace that runs a private, on-device AI model with zero internet required — and lets you jump to the best cloud models from the very same conversation, on the very same account. No separate app for offline mode, no syncing chats between two products.

OpenAIGPT-4o
ClaudeClaude 3.5
GeminiGemini
GrokGrok 2
OpenAIGPT-4 Turbo
ClaudeClaude Opus
GeminiGemini Ultra
OpenAIGPT-4o
ClaudeClaude 3.5
GeminiGemini
GrokGrok 2
OpenAIGPT-4 Turbo
ClaudeClaude Opus
GeminiGemini Ultra

Fully offline

No internet, no server round-trip. The model runs on your phone's or laptop's own CPU.

Private by construction

Nothing leaves the device — not the prompt, not the reply. There's no network call to intercept.

Cloud when you want it

The same account, same chat window, also reaches Claude, GPT, Gemini and Grok whenever you're online.

Free once unlocked

Local inference runs on your hardware, not ours — it never draws from your credit balance.

Which models actually run on-device

No black box — here is the exact catalog we ship today, with real download sizes and RAM requirements, not a marketing round number. Every model is open-weight under a commercial-use license (Apache-2.0, MIT or Llama-Community).

Android — local LLM app
Phone-sized models, 2–4 GB RAM
ModelLicenseSizeMin RAMContext
Qwen 2.5 1.5B
Alibaba
Apache-2.0~940 MB2 GB8K
LFM2.5 1.2B
Liquid AI
LFM-1.0~697 MB3 GB32K
Llama 3.2 3B
Meta
Llama-Community~1.9 GB4 GB4K
Llama 3.2 1B
Meta
Llama-Community~770 MB2 GB4K
Windows desktop — local LLM app
Larger models, 8–16 GB RAM
ModelLicenseSizeMin RAMContext
Qwen 2.5 7B
Alibaba
Apache-2.0~4.7 GB8 GB32K
Mistral 7B Instruct v0.3
Mistral AI
Apache-2.0~4.4 GB8 GB32K
Llama 3.1 8B
Meta
Llama-Community~4.9 GB8 GB128K
Qwen 2.5 Coder 7B
Alibaba
Apache-2.0~4.7 GB8 GB32K
Phi-4 14B
Microsoft
MIT~9.1 GB16 GB16K

The hardware reality: we don't publish a single tokens-per-second number, because one figure can't represent every phone or laptop's CPU — that would be marketing, not fact. What we can tell you: if your device meets the minimum RAM for a model, it runs. Mobile models (1–3B, quantized to Q4_K_M) typically answer within a couple of seconds on a modern phone; the larger 7–14B desktop models take longer per response but run comfortably on an 8–16 GB laptop with no GPU required. Older or lower-RAM hardware will be slower — that's the honest trade-off of running fully offline with no cloud GPU behind it.

Local LLM + cloud, in one app

Assign a local model to one assistant for fully offline, private chat — nothing leaves your device. Assign Claude, GPT, Gemini or Grok to another for cloud-scale reasoning, web search or image generation. Both assistants live in the same workspace, under the same account, and you switch between them mid-project without exporting or re-importing a single message.

1

Download a model

Pick one from the catalog above, right on your phone or desktop app — no dev tools, no GitHub clone.

2

Assign it to an assistant

The assistant now runs entirely on-device. Chat with it in airplane mode, on a plane or with no signal at all.

3

Switch to cloud anytime

Open another assistant on Claude, GPT, Gemini or Grok in the same workspace when you want more power.

Available from the Light plan up

Local LLM is included on Light ($5/mo), Pro, Pro Plus, and every Team plan. The Free plan doesn't include it. Once unlocked, running a local model costs nothing per message — it's your device's CPU doing the work, not ours, so it never touches your credit balance.

Compare plans

Get it on your device

Local LLM ships today on Android and the Windows desktop app. macOS and iOS are coming — see the full download page for current availability on every platform.

Windows
Windows 10 / 11
Download
Android
Phones & tablets
Google Play

Try offline AI free — no card needed

Sign up, download the app, and load a local model in a couple of minutes. Switch to cloud models whenever you want more power.

Frequently asked questions

Can I use AiMixup AI chat without an internet connection?

Yes. Assign a local model to an assistant on Android or the Windows desktop app and the whole conversation runs on your device — no network call, no cloud request. Turn on airplane mode and it keeps working.

Is there a local LLM app for Android?

Yes, the AiMixup Android app ships four on-device models (Qwen 2.5 1.5B, LFM2.5 1.2B, Llama 3.2 3B, Llama 3.2 1B). Download one from Settings → Local models and assign it to an assistant to chat fully offline.

Can I run an LLM on my phone without a subscription to a cloud AI?

Once a model is downloaded, running it costs nothing per message — inference happens on your phone's own CPU, not on our servers, so it never touches your credit balance. A Light plan or above is required to unlock the feature; the model itself is then free to use as much as you want.

Is local LLM available on iOS?

Not yet — today AiMixup's on-device engine runs on Android and the Windows desktop app. iOS and macOS support is planned; see /download for current app availability.

How much RAM do I need to run an AI model locally?

It depends on the model. Phone models start at 2 GB RAM (Llama 3.2 1B, Qwen 2.5 1.5B) up to 4 GB (Llama 3.2 3B). Desktop models need more: 8 GB for the 7–8B models (Qwen 2.5 7B, Mistral 7B, Llama 3.1 8B) and 16 GB for Phi-4 14B, the highest-quality option we ship.

Is my chat private when I use a local model?

Yes — that's the point of running it locally. The prompt, the model weights and the generated reply all stay on your device; nothing is sent to AiMixup or to any AI provider. It's the strongest privacy option we offer, stronger than any cloud model's data policy, because there is no network call to make.

Can I use local and cloud AI models in the same app?

Yes — that's the combination nobody else ships. Create one assistant on a local model for fully offline, private chat, and another on Claude, GPT, Gemini or Grok for cloud-scale reasoning, image generation or web search — both live in the same workspace, same account, same subscription.

Which plan do I need for local LLM?

Light and above (Pro, Pro Plus), and every Team plan. The Free plan doesn't include it. Local inference itself doesn't consume credits once unlocked — it runs on your hardware.

Which AI models can I download to run locally?

On mobile: Qwen 2.5 1.5B, LFM2.5 1.2B, Llama 3.2 3B and Llama 3.2 1B. On Windows desktop: Qwen 2.5 7B, Mistral 7B Instruct v0.3, Llama 3.1 8B, Qwen 2.5 Coder 7B and Phi-4 14B. All are open-weight models under commercial-use-permitted licenses (Apache-2.0, MIT or Llama-Community).

How fast is a local model compared to a cloud model like GPT or Claude?

It depends entirely on your device's CPU — we don't publish a single tokens-per-second number because one figure can't represent every phone or laptop. In general, the 1–3B mobile models respond within a couple of seconds per message on a modern phone; the larger 7–14B desktop models take longer per response but run comfortably on an 8–16 GB laptop. Slower or older hardware will be slower; that trade-off is the cost of running fully offline with no cloud GPU behind it.