최근 관측된 다운로드 가격은 $17.99입니다. 현재 결제 가격은 공식 App Store 링크에서 확인하세요.
현재 선택한 스토어에서 관측된 다운로드 가격입니다. 지역별 가격에서 국가별 환산 가격을 비교할 수 있습니다.
Apple 공개 App Store 페이지에서 현재 확인되는 항목입니다. App Store Connect의 전체 상품 목록과 다를 수 있습니다.
출처: Apple 공개 App Store 페이지 · 공개 항목만 포함
Your Phone Is Now an AI Server
LLM Server turns your iPhone or iPad into a private AI inference server. Run large language models entirely on-device, expose an OpenAI-compatible API to your local network, and chat with your model from any browser on any device — laptop, desktop, tablet, even another phone. No cloud, no subscription, no data leaving your hardware.
Chat From Any Browser on Your Network
Start the server and open the URL on any device sharing your Wi-Fi. The built-in web interface gives you a clean chat UI instantly — no client app to install, no account to create. Your laptop, your partner's tablet, a colleague's machine: all talking to a model running on your phone.
OpenAI-Compatible API, Drop-In Ready
A standard OpenAI-compatible API means LLM Server slots into any tool you already use — Continue, Open WebUI, LangChain, custom scripts, the lot. Chat completions, text completions, streaming (SSE), and model listing endpoints all supported. Ollama CLI commands work too.
Fully Offline, Fully Yours
- Download a GGUF model once and you're done with the cloud. Inference runs locally on Apple Metal GPU with no internet required. Your prompts, your conversations, your data — none of it leaves the device. Airplane mode works fine.
Any GGUF Model From Hugging Face
- Browse and download directly from Hugging Face with built-in search, or import your own files. LLaMA, Mistral, Phi, Gemma, Qwen, DeepSeek, and every other llama.cpp-supported architecture runs out of the box. Background downloads with progress tracking so you can keep working.
Enterprise-Grade Security
- TLS/HTTPS encryption — generate self-signed certificates or import your own chain and private key.
- API key authentication — Bearer tokens with per-key management. Generate cryptographically secure keys or bring your own.
- Bind control — lock to localhost, open to your LAN, or pin to a specific interface.
Tune Every Knob
Full control over generation: context size up to 32K, temperature, top-p, top-k, repeat penalty, frequency and presence penalties, max tokens, seed, GPU layer offloading, and thread count. Save presets globally or per model.
Smart Resource Management
Your phone stays responsive under load. Real-time thermal monitoring with automatic thread reduction under pressure and request rejection at critical temperatures. Memory-aware model loading with conservative budgeting. Configurable request queues and per-request timeouts.
Built for Developers
- Live API docs with copy-paste curl examples
- Structured logging (debug, info, warning, error)
- One-tap copy for server addresses and API keys
What's Inside
Dashboard with one-tap server control, model manager with download progress, complete settings hub (server, inference, security, API keys, developer tools), and guided onboarding for first-time setup.
최근 관측된 App Store 데이터를 바탕으로 한 빠른 답변입니다.
최근 관측된 다운로드 가격은 $17.99입니다. 현재 결제 가격은 공식 App Store 링크에서 확인하세요.
Apple 공개 페이지에는 현재 이 스토어에서 0개의 구매 항목이 표시됩니다. 공개 목록은 완전하지 않을 수 있습니다.
네. App Store 가격은 스토어, 통화, 세금 및 개발자 가격 설정에 따라 달라질 수 있습니다. 현재 1개 지역의 관측 데이터가 있습니다.