Cnuas
Cnuas (pronounced Knoo-us, the Irish word for a cluster) is an experimental, open-source AI/HPC rack-scale emulation platform under development. Its architecture is based on the Open Compute Project Open Rack v3 specifications, bringing device, driver, fabric and rack-management interfaces into a software environment for research and development. It is both a software product and an extensible framework for community collaboration.
Cnuas is under
active research and release preparation. The paper is published as an arXiv preprint, and the
six-patch QEMU v2 contribution series has been submitted to qemu-devel. Source
repository publication is pending upstream merge of the QEMU and Linux patches.
Research and availability status.
Virtual GPUs, virtual GPU-peer fabric, virtual InfiniBand and RoCE switches and virtual RDMA NICs, all built from scratch and developed the way the Linux kernel is, so unmodified Linux guests run distributed workloads on emulated silicon.
The Cnuas paper is available as a preprint on arXiv. If you are an academic user, please cite Cnuas using the BibTeX entry below. The other styles are generated from that same entry.
The paper will be presented at a conference. This page will carry the proceedings reference once it is available.
@misc{janjua2026cnuassoftwaredefinedaihpcrackscale,
title={Cnuas: A Software-Defined AI/HPC Rack-scale Emulation Platform and Hyperscale Data Center Facility Twin},
author={Weqaar Janjua and Eoin OConnell and Mihai Penica},
year={2026},
eprint={2609.15889},
archivePrefix={arXiv},
primaryClass={cs.DC},
url={https://arxiv.org/abs/2609.15889},
}
Janjua, W., OConnell, E., & Penica, M. (2026). Cnuas: A Software-Defined AI/HPC Rack-scale Emulation Platform and Hyperscale Data Center Facility Twin (arXiv:2609.15889). arXiv. https://doi.org/10.48550/arXiv.2609.15889
Janjua W, OConnell E, Penica M. Cnuas: A Software-Defined AI/HPC Rack-scale Emulation Platform and Hyperscale Data Center Facility Twin [Internet]. arXiv; 2026. Available from: https://arxiv.org/abs/2609.15889 doi:10.48550/arXiv.2609.15889
Janjua, W., OConnell, E. and Penica, M. (2026) “Cnuas: A Software-Defined AI/HPC Rack-scale Emulation Platform and Hyperscale Data Center Facility Twin.” arXiv. Available at: https://doi.org/10.48550/arXiv.2609.15889.
cnuas is the unified control plane. A single command surface drives the switch
fabric, the GPU fabric, GPUs, RDMA NICs and VM lifecycle, and the same service layer is
exposed as a REST API for automation. Every CLI capability is also an API endpoint.
$ cnuas system health ┏━━━━━━━━━━━━━┳━━━━━━━━━━━┓ ┃ Component ┃ Available ┃ ┡━━━━━━━━━━━━━╇━━━━━━━━━━━┩ │ cnuasswitch │ yes │ │ cnuaslink │ yes │ │ cnuasgpu │ yes │ │ cnuasnic │ yes │ │ vm │ yes │ └─────────────┴───────────┘ $ cnuas switch set-mode 2 STRICT_IB { "status": "ok", "port": 2, "mode": "STRICT_IB" } $ cnuas vm up vm-a --gpus 1 vm-a running, CnuasNIC + CnuasGPU attached
Cnuas does not use host-based shortcuts like Linux bridges or Soft-RoCE. Instead it builds custom QEMU PCIe devices and their corresponding Linux kernel drivers, so the guest sees a real RDMA HCA and a real GPU on its PCIe bus. From the application point of view it is talking to hardware, even though every byte is moving through software.
cnuas_net.ko, cnuas_ib.ko) probe the device and register with ib_corelibibverbs works unmodified, so ibv_rc_pingpong, perftest and NCCL all runNot all developers, researchers and students can get access to a rack-scale GPU cluster, or to the cutting-edge AI/HPC GPU and accelerator platforms inside it. Existing Linux virt approaches (bridges, Soft-RoCE) hide the hardware semantics that matter, such as DCB, ECN marking, IB LIDs and GPU peer DMA. Cnuas keeps those semantics in software so what you learn here transfers directly to production-grade AI/HPC GPU and accelerator platforms.
Cnuas is built from four hardware components, each its own open-source project under PacketFive, plus a software stack that mirrors the NVIDIA developer experience. Cnuas ties them together as the integration and orchestration layer.
PCIe device model exposing standard BARs, doorbells and DMA queues. The companion
kernel drivers register a netdev and an RDMA HCA. Guest applications use libibverbs
as if it were Mellanox or NVIDIA hardware.
Standalone C daemon acting as a top-of-rack switch, running on the host with no guest. An
epoll event loop listens on
per-port UNIX sockets, classifies frames (Ethernet, RoCEv2 or InfiniBand LRH), and
forwards using a per-pipeline forwarding table over a JSON-over-UDS control plane.
SIMT accelerator with Streaming Multiprocessors, warp scheduling, shared memory and
HBM-equivalent device memory, with the tensor MAC applied inside the matrix multiply
rather than exposed as a separately addressable tile unit. It runs in three forms that
execute the same kernels: a Soft-GPU inside an ordinary host process, with no
hypervisor, guest or kernel module; a host character device, where
cnuasgpu_host.ko presents /dev/cnuasgpu_hostN without any PCI
device present; and a QEMU PCIe device exposing /dev/cnuasgpuN for
guest-side and rack deployments. It is building toward a
complete NVIDIA-equivalent
userland. Today that means Cnuas Compute, the CUDA-equivalent runtime, cnuassmi, and a
cnuascc compiler that lowers device source to CnuasIR for the scalar subset; vector
code generation and CnuasCCL collectives are in progress.
Independent GPU-to-GPU interconnect, an NVSwitch equivalent, served by a host daemon that
needs no guest. A VM with a CnuasGPU but no
CnuasNIC can still do GPU peer DMA. It uses ivshmem as the fast path for
cross-VM shared memory. This is GPU to GPU. Getting a NIC to read GPU memory directly is
a separate path, Cnuas peer memory,
which is experimental.
CnuasGPU is building toward a complete NVIDIA-equivalent userland, so the same programming model and tooling students use here transfer directly to production GPUs. Each tool uses clean Cnuas naming and mirrors a familiar NVIDIA counterpart. Components marked Planned are specified in the design documents but carry no code in the tree yet, Partial means some parts ship and others are specified only, Experimental means the code is written and its unit tests pass but the end-to-end result is not yet claimed, and every Shipping claim is backed by a named test in the Validation Matrix. See the Software Stack Datasheet for the full breakdown.
The CUDA-equivalent runtime stack. libcnuasrt.so mirrors the CUDA
Runtime and libcnuasdev.so the device API.
Device compiler that lowers GPU kernels to CnuasIR, the CnuasGPU intermediate representation. The front end and scalar code generation are implemented, so a kernel now compiles to an object; vector code generation and the Clang and LLVM pipeline are still to come.
nvcc + PTX PartialMulti-GPU collective communications, including AllReduce and Broadcast, running over the CnuasLink peer fabric with a RoCE fallback. v0.1 runs over the bootstrap channel.
NCCL PartialOne-sided GPU-to-GPU shared memory for partitioned workloads, using CnuasLink put and get as the transport.
NVSHMEM PlannedDevice management and monitoring CLI. It lists GPUs, reports utilisation and memory, and exposes realistic telemetry.
nvidia-smi ShippingFleet health checks and diagnostics, reading per-SM and chip-wide counters for cluster management.
DCGM PlannedKernel profiler that surfaces per-SM occupancy and chip-wide performance counters for tuning workloads.
Nsight Compute PlannedCnuasBLAS v0.1 ships as libcnuasblas.so, providing dense linear algebra over the
runtime kernels. CnuasDNN, CnuasFFT, CnuasSPARSE and CnuasSOLVER are specified only.
Lets a CnuasNIC memory region point straight at a CnuasGPU device-memory allocation, over
the upstream Linux DMA-BUF RDMA interface rather than the removed out-of-tree peer-memory
client API. The exporter, libcnuaspeermem, the standard
ibv_reg_dmabuf_mr registration path and the NIC importer are implemented and
the ABI and ownership tests pass. Live transfer validation still needs a guest exposing
both devices, so end-to-end operation is not claimed yet.
Exposes CnuasGPU devices to Docker and other container runtimes for reproducible GPU workloads.
nvidia-container-toolkit PlannedCnuas is a combination of implemented prototypes, work in progress and planned development. The labels below are used through the rest of this page, so a heading says which of the three it belongs to.
Implemented components with bounded prototype evidence: the switch fabric, the RDMA NICs, the kernel modules, the control plane and the rack-management firmware. Broader integration and evaluation continue.
An early research implementation. The compiler front end, the IR, the collectives and the math libraries each carry their own status, and further language, runtime, workload and interoperability development is planned.
An early experimental Cnuas extension under research and development. It is not an equally mature core capability, and its outputs are unsuitable for engineering, procurement, safety or operational decisions.
The source is being released in stages. Implemented code, demonstrated integration, physical validation and general readiness are separate questions, and the research and availability status page records where each one stands.
Day-to-day work sits on the public GitHub project board, so what is in progress, what is blocked and what is queued can be read directly rather than inferred from this page.
Configure, set up and interact with the entire system from a scriptable CLI or a REST API. Both front ends call the same service layer, so nothing is CLI-only or API-only.
Interactive Rich tables by default, and --json on any
command for scripting. Command groups cover every domain of the rack.
cnuas switch, RoCE and InfiniBand fabric, 18 commandscnuas fabric, CnuasLink GPU peer fabriccnuas gpu, CnuasGPU device controlcnuas nic, CnuasNIC and RDMA device controlcnuas vm, virtual-machine lifecyclecnuas system, inventory, health and versionsA FastAPI service (cnuas api) with OpenAPI docs, so any
language or orchestrator can drive the platform over HTTP.
$ cnuas api --host 0.0.0.0 --port 8080 # OpenAPI docs at http://host:8080/docs $ curl host:8080/api/v1/switch/ports [{ "port": 0, "mode": "AUTO", "link": "up" }, { "port": 2, "mode": "STRICT_IB", "link": "up" }]
The reference deployment is based on the Open Compute Project Open Rack v3 specifications, the rack design the modern AI buildout was created around. Two 1OU TOR switches (CnuasSwitch and CnuasLink) sit at the top of the rack, above a management host and eight 2OU VM blades acting as GPU compute nodes. Below them the rack carries what an ORv3 rack carries: a pair of 1OU power shelves feeding the 48 V busbar, a 3OU battery backup shelf at the foot, and blanking panels over the bays this deployment does not use. The U-heights are visual in the emulator, and the layout follows the specification, which is what makes the same diagrams useful for teaching, design and lab build-outs.
The CnuasSwitch fabric carries every Ethernet frame, including RoCEv2 and RDMA
on UDP 4791, the same path NCCL's IB transport, perftest and
ibv_rc_pingpong flow over. The CnuasLink fabric is independent
and carries only GPU-to-GPU peer traffic. A guest can use one, the other, or both,
exactly as in a real GPU rack.
cnuas-webui is a client of the same management socket the CLI uses, so it adds no
switching logic of its own. Its rack and switch views give a developer visual context for
where equipment sits, how it is interconnected and which switch ports are up, which is the
part of rack-scale work that software interfaces normally keep out of sight. The intent is to
let someone relate a workload to the infrastructure underneath it. Whether that changes
learning outcomes has not been tested.
This is separate from the exploratory OpenUSD facility extension further down the page. The web interface drives the emulated rack. The facility extension generates campus scenes.
The Cnuas Facility Twin is an early experimental extension exploring hyperscale data-center modeling with OpenUSD. The current prototype generates illustrative campus scenes, performs simplified load calculations and connects readings from emulated rack power equipment to scene attributes. NVIDIA Omniverse and Isaac Sim are intended exploration environments. Further development will focus on scenario validation, reference data, calibration and reliable integration, with community collaboration alongside the Cnuas core.
Those are illustrative inputs and the arithmetic that follows from them, not measured capability or validated capacity. The walkthrough film, the rack level data hall plan, the full load roll-up, the commands that regenerate all of it and the research roadmap are on the simulation page.
The campus is one input file. The planner reads a JSON campus specification, so you can describe your own site, buildings, halls, rack rows and plant, have the configuration and the spatial arithmetic checked, and export a stage of your own. Prototype outputs are unsuitable for engineering, procurement, safety or operational decisions.
The design documents and datasheets describe what Cnuas is. The guides describe how to install it, how to run it and how to extend it, and the roadmap records what is built, what is planned and what is blocked. Day-to-day work is tracked in public on the GitHub project board, so the roadmap tables, the project plan graphic and the open issues are the same set of items. Three of the planned workstreams are summarised below the guides.
Size a host before installing. Cnuas has a host-only form that needs no virtualisation, no root and no guest kernel, and knowing which form you want saves most of the setup.
Requirements →Install the host-only or the rack-scale software, bring up QEMU, the ORV3 shelf and the datacentre simulation.
Install →Configure and operate Cnuas day to day: environment, CLI, REST, Redfish, the power shelf and the facility twin.
User guide →Build Cnuas with Bazel or Make, use the C and Python APIs, and meet the coding, testing and CI standards the tree is held to.
Developer guide →Add a new emulated component, either on the ORV3 RS-485 segment with its Modbus register map, or reached over Ethernet and IP.
Extend →Every workstream, its status and the open design decisions behind it. Epics and Tasks on the board link straight to the issues, pull requests and commits in the component repos. The project plan draws the same board as a roadmap graphic you can pan, zoom and download as a PDF.
Roadmap → Project plan → Project board ↗Host and device coherent-memory interconnect semantics, assessed against the target CXL specification and what QEMU and Linux already provide, then prototyped to a scoped guest-visible topology. It comes first in the work order.
Accelerator-peer attachment, taken up after the CXL research. The order is a sequencing decision rather than a technical dependency, and both are distinct from the CnuasLink fabric that exists today.
Berkeley UCCL is an upstream research collective library, separate from the implemented CnuasCCL. The open tasks cover compatibility assessment, bounded host-memory P2P experiments and AllReduce correctness with reproducible evidence. No GPU interoperability is claimed, and no dates are set.
P2P task, CnuasNIC #25 ↗ AllReduce task, CnuasNIC #26 ↗Every Cnuas component, hardware, software and firmware, has a product datasheet. Each one is generated from a single Markdown source into three editions, so the web page, the printable PDF and the plain-text file cannot drift apart. Document ids, revisions and statuses below are read from the datasheet build itself. Published means implemented and covered by tests, Partial that some parts ship and others are specified only, Experimental that the work is a research preview or that the code is written and its unit tests pass while the end-to-end result is still gated, and Preview that the component is designed but carries no implementation yet.
| Document | Datasheet | Part | Status | Download |
|---|---|---|---|---|
| Hardware | ||||
DS-CNU-001 rev D |
Cnuas Platform | Platform |
Published | PDF · TXT |
DS-CNU-002 rev C |
CnuasSwitch | cnuas-vswitchd |
Published | PDF · TXT |
DS-CNU-003 rev D |
CnuasNIC | cnuas-vnic |
Published | PDF · TXT |
DS-CNU-004 rev L |
CnuasGPU | cnuasgpu |
Published | PDF · TXT |
DS-CNU-005 rev C |
CnuasLink | cnuasgpu-link-switchd |
Published | PDF · TXT |
| Software | ||||
DS-CNU-010 rev E |
Cnuas Software Stack | Software map |
Published | PDF · TXT |
DS-CNU-011 rev D |
CnuasRT | libcnuasrt.so |
Published | PDF · TXT |
DS-CNU-012 rev K |
CnuasDev | libcnuasdev.so |
Published | PDF · TXT |
DS-CNU-013 rev D |
Kernel Modules | cnuas_net.ko, cnuas_ib.ko, cnuasgpu.ko, cnuasgpu_host.ko |
Published | PDF · TXT |
DS-CNU-014 rev B |
Verbs Provider | libcnuas-rdmav34.so |
Published | PDF · TXT |
DS-CNU-019 rev B |
Math Libraries | libcnuasblas.so and four specified |
Partial | PDF · TXT |
DS-CNU-020 rev A |
Control Plane | cnuas, cnuas-api |
Published | PDF · TXT |
DS-CNU-021 rev A |
Cnuas Tools | cnuas-tools |
Published | PDF · TXT |
DS-CNU-022 rev A |
Management Tools | cnuassmi, cnuas-cli, cnuaslink-cli |
Published | PDF · TXT |
DS-CNU-023 rev A |
Cnuas Peer Memory | libcnuaspeermem |
Experimental | PDF · TXT |
DS-CNU-032 rev B |
Calibrated Virtual Time | cnuas-timing |
Published | PDF · TXT |
| Firmware and facility | ||||
DS-CNU-031 rev C |
CnuasBMC | cnuas-rackmond |
Published | PDF · TXT |
DS-CNU-030 rev C |
Facility Twin | cnuas-facility |
Experimental | PDF · TXT |
| Preview stack | ||||
DS-CNU-015 rev C |
CnuasCC | models nvcc |
Partial | PDF · TXT |
DS-CNU-016 rev D |
CnuasIR | models PTX |
Partial | PDF · TXT |
DS-CNU-017 rev D |
CnuasCCL | models NCCL |
Partial | PDF · TXT |
DS-CNU-018 rev A |
CnuasSHMEM | models NVSHMEM |
Preview | PDF · TXT |
Cnuas is an open platform for infrastructure research and hands-on learning, and it is the environment PacketFive Academy courses run against. Students and researchers run real distributed-training workloads, RDMA microbenchmarks and MPI collectives against the emulated fabric, building the skills modern AI cluster operations demand, on a laptop.
A physical testbed gives you real hardware and real timing, which nothing here replaces. What it usually does not give you is the accelerator itself. The installed parts are vendor products, so you can change the application and the system software around them but not their RTL, firmware, math blocks or fabric logic. Cnuas opens that surface. The device model and experimental RTL, the ISA, the compiler, the drivers, the runtime, the collectives, the network, the rack-management firmware and the facility model are all readable and all modifiable, so a cross-layer experiment can change both sides of an interface and still be validated on one workstation.
No physical GPUs, NICs or switches required. A multi-node GPU cluster fits on a single workstation.
IB LIDs, GRH and LRH headers, DCB pause frames and ECN marking, all the things that matter for real production debugging.
Open-source under the Apache License 2.0, modifiable and reproducible. Ideal for academic networking and HPC research where commercial silicon is a black box.
Cnuas stands on the open-source stack the industry runs on. Compute and devices are
emulated with QEMU, the rack architecture is based on the Open Compute Project Open Rack v3
specifications, and the out-of-band management plane integrates OpenBMC and Renode for realistic
Baseboard
Management Controller behaviour, serving the DMTF
Redfish
1.17.0 API so real management clients drive the rack unmodified. The campus around
that rack is written as an
OpenUSD
stage by the experimental Facility Twin extension, so it opens in NVIDIA Isaac Sim,
usdview, Blender or any other
USD-capable tool. Beyond emulation, CnuasGPU targets the Microchip PolarFire SoC FPGA on
the Icicle Kit, a PCIe-attached board with hard RISC-V cores, so the same accelerator
stack can be validated on real silicon. The RTL for that fabric is in the tree and an
eight-lane FP32 descriptor engine passes functional simulation against host-computed
references. That is a functional result, not a timing-accurate one, and no bitstream, timing
closure, resource report or board validation is claimed. The FPGA attachment is active
research work, not a released feature.
Device and machine emulation
Open Rack v3 specifications
Redfish 1.17.0, the DMTF standard the
BMC serves for out-of-band management
Campus written as an OpenUSD stage,
with Isaac Sim as the intended exploration environment
QEMU, Open Compute Project, OpenBMC, Renode, OpenUSD, NVIDIA, Isaac Sim, Microchip and PolarFire, and their logos, are trademarks of their respective owners. They are shown to indicate the technologies Cnuas builds on and interoperates with. Cnuas is not affiliated with, sponsored by or endorsed by these projects or companies. The "Built on OpenBMC" logo is by the OpenBMC Project, used under CC BY 4.0. The NVIDIA logo is reproduced unaltered from NVIDIA's logo and brand usage page. OpenUSD was created by Pixar and is governed by the Alliance for OpenUSD, of which NVIDIA is a founding member; it is grouped above because Cnuas uses it together with Isaac Sim, not because NVIDIA owns it.
Cnuas is a product of PacketFive, the trading name of Packet Five Networks Limited, a research and development company based in Dublin, Ireland. We build open infrastructure for the AI/HPC datacenter, spanning simulation, training and facility security.
Cnuas is developed in the open and is the environment PacketFive Academy training courses run against. For commercial enquiries, research collaboration or support, get in touch with the team.
Packet Five Networks Limited, 51 Bracken Road, Sandyford, Dublin D18 CV48, Ireland. CRO 764119.