Frontier open-weight LLMs, hosted on the NRP

Free access for researchers and educators to a rotating catalog of frontier open-weight models — through a hosted chat interface, ready-made coding-CLI configs, and an OpenAI-compatible API with no per-token billing, from single chats to batch runs over whole corpora.

Open WebUI chat running on the NRP

What you can use them for

The same endpoint carries a chat window, a coding agent, the platform's own operations, and corpus-scale batch jobs — none of it billed by the token.

Document understanding and sentiment

Researchers can parse and summarize large document collections, then extract signals such as sentiment from sources like securities filings.

Agentic coding for classrooms and research groups

Classrooms and research groups can use coding agents with NRP-hosted models to develop, review, and iterate on software projects.

Batch LLM processing

No per-token billing and no cap on how many requests you send: point a script at the API and work through whole document corpora, eval sweeps, or synthetic-data jobs. Fair use governs how many requests run at once — never how many you make.

Agentic operations

The platform turns its own operations over to the models it hosts: operators question failing nodes, users untangle problem pods, and the accounting system answers in plain language.

connected surfaces
nodes
Query and diagnose failing infrastructure in the context of the cluster it runs.
pods
Inspect logs, events, and related pod details to troubleshoot a workload.
accounting
An MCP server connects models to the NRP accounting system to surface usage insights.

How to get access

Step 1: Get an NRP account

Sign in once with your institutional or research identity to register a Nautilus account.

Step 2: Request an LLM-enabled namespace

Reach out and we will enable LLM access on your namespace so its members can mint API keys.

Step 3: Generate your LLM API key

Use the LLM API Keys page to create a personal API key, then plug it into Open WebUI, Chatbox, or any OpenAI-compatible client.

Models we host

Frontier open-weight models with strong reasoning, coding, and multimodal capabilities. Pick the one that fits your task.

qwen3

180B (6B active) · 1M ctx

Flagship frontier multimodal MoE — a Qwen4-architecture preview, and the catalog's strongest agentic coder.

qwen3-small

27B · 1M ctx

Compact Qwen3.8 — multimodal, agentic, low-latency.

gpt-oss

120B · 131K ctx

OpenAI's open-weights agentic model — tiny GPU footprint, strong tools, LTS candidate.

gemma

31B · 262K ctx

Google's Gemma 4 — multimodal, efficient frontier performance.

gemma-small

12B · 262K ctx

evaluating

Tiny Gemma 4 with unique audio input — ASR and speech-to-text on a 12B model.

kimi

1T · 131K ctx

evaluating

Moonshot's 1T-parameter frontier coding model with multimodal inputs.

glm-5

753B · 1M ctx

evaluating

Z.ai's 753B frontier coding model with NVFP4 weights — top of the catalog on Artificial Analysis' Intelligence Index.

deepseek-v4-flash

304B · 1M ctx

evaluating

DeepSeek's 304B MoE, served from the experimental vision checkpoint — image input with a full million-token context window.

minimax-m2

230B · 205K ctx

evaluating

Efficient frontier coding model — 230B in native FP8, fits comfortably on four A100s.

qwen3-embedding

8B · 262K ctx

Multimodal embedding model for retrieval and vector search — not a chat model.

Use the client you already have

Hosted chat interfaces, desktop apps, and coding CLIs all work with NRP-hosted LLMs through the OpenAI-compatible endpoint.

Open WebUI
Chatbox
Cherry Studio
Claude Code
OpenCode
Crush
Kimi CLI
Copilot CLI
endpoint
https://ellm.nrp-nautilus.io/v1

Requests authenticate with a personal API key as a bearer token — mint one on the LLM API Keys page once your namespace has LLM access enabled.

Ready to get started?

We will get you onto NRP-hosted LLMs and answer any questions about access, namespaces, and API keys.