Skip to content

Joining the NRP

Contributing hardware  —  For institutions and campus IT staff adding servers to the NRP pool.

Contributing a node to the NRP is a trade. You supply the hardware, the rack space, and the network path. The NRP installs Ubuntu Server 24.04 as the base operating system, joins the node to the Nautilus cluster, and runs it from there — patching, monitoring, and upgrades included. In return your researchers get a share of the entire pool, not only the machine you bought.

You keep priority on your own hardware. When your people want the nodes you contributed, they get them. When those nodes are idle, the rest of the community runs opportunistic work on them. That is the whole mechanism: nothing sits unused, and every contributor draws on far more capacity than they put in.

The NRP is funded by the National Science Foundation rather than cost recovery, so there is no fee to participate and no charge for your researchers’ access.

Who can join

Using the platform is open to U.S. nonprofit research and education institutions. Community colleges qualify, not only R1 universities.

Contributing hardware comes from that same community: U.S. researchers and the collaborators they work with can add nodes to the pool.

What your researchers get

  • The whole GPU pool, not just your nodes — plus FPGAs and other specialized hardware.
  • Hosted JupyterHub and Coder, so people who will never write a Kubernetes manifest can still use the cluster. A private JupyterHub is an option for groups that want one.
  • Direct cluster access with kubectl against the Nautilus API, for those who do want it.
  • Petabyte-scale S3 and Ceph storage shared across the cluster.
  • A hosted LLM endpoint — OpenAI-compatible, with per-user API keys.
  • Training and support — recorded Docker, Kubernetes, and JupyterHub sessions, bi-weekly office hours, and the Nautilus Support chat.

Responsibilities

The split is deliberately lopsided. You own the physical plant and the network path; the NRP owns everything from the firmware up.

Contributing site

Required

  • Network connectivity A Science DMZ path to the node at 10–100G with jumbo frames, outside the campus firewall.
  • Management access ACL Permit the NRP management address to reach the node's IPMI interface.
  • Rack space and cabling Rack units, power drops, and network cabling for the node.
  • Power and cooling Conditioned power and environmental control the node can stay up on.
  • On-site contact A named person we can reach for physical access, reboots, and hardware swaps.
  • Security incident notice Tell us about incidents at your site that could affect the node you contribute.

Suggested

  • Warranty coordination Forward vendor hardware tickets to the NRP for triage.
  • Maintenance notice Tell us before planned power or network downtime so we can drain the node.

NRP

Required

  • Firmware and OS BIOS updates, Ubuntu Server 24.04 installation, and lifecycle management.
  • Security from the OS up Patching, package updates, and host hardening.
  • Monitoring and response 24/7 alerting, health checks, and incident handling.
  • Security incident notice We tell you about incidents on the cluster that could affect your site.
  • Cluster network policy Central Calico policy firewalls the node, so you maintain no per-node rules.

Suggested

  • Procurement guidance Hardware recommendations and fulfillment support.
  • Onboarding Cluster join, storage provisioning, and GPU drivers.
  • Capacity reporting Grafana dashboards and accounting for your own utilization.
  • Docs and training Reference material and sessions for your local administrators.

Two of these deserve the detail on the Networking page before you commit: the node needs a Science DMZ path with 9000 MTU to every other node in the cluster, and local iptables or firewalld has to be off so the cluster-wide Calico policy can manage the host firewall centrally.

Security

The NRP monitors the security of the nodes it runs and applies security patches to their software continuously. You do not track CVEs, schedule OS updates, or maintain host firewall rules for a contributed node — that work is ours from the firmware up.

Incident disclosure runs both ways:

  • You tell us about security incidents at your site that could affect the node you contribute — a compromised management network, credentials exposed, anything that puts the host at risk.
  • We tell you about security incidents on the cluster that could affect your site.

Report either direction through Nautilus Support.

What to send us

  1. Node IP configuration — public IP, subnet mask, and gateway.

  2. DNS servers for the host network

    If you have no preference we default to public resolvers: 1.1.1.1 (Cloudflare) and 8.8.8.8 (Google).

  3. IPMI access — the address, and credentials if any are set. Include jump-host details if the interface is not directly reachable.

  4. Firewall allowance for IPMI — if your ACL needs an entry, the NRP management address is 67.58.53.149.

  5. Confirmation you have read the network requirements

That is enough for us to build the host.

Submit the details

Send the list above through Nautilus Support, the Matrix chat that is our primary support channel. We take it from there: configuring the node, joining it to Nautilus, and bringing it into the scheduling pool.

Questions before you commit are welcome at the same address, or at the bi-weekly office hours.

After the node is online

Your researchers still need accounts, and that path is separate from the hardware. It runs through CILogon federated identity, a guest-to-user promotion by an existing admin, assignment to a namespace, and accepting the Acceptable Use Policy. Namespace admins are personally responsible for activity in their namespaces.

Get Access is the full walkthrough — point your researchers there rather than repeating it locally.

U.S. National Science Foundation

This work was supported in part by National Science Foundation (NSF) awards CNS-1730158, ACI-1540112, ACI-1541349, OAC-1826967, OAC-2112167, CNS-2100237, CNS-2120019.