Full-liquid W45 direct-to-chip (fanless) · two-loop · 3× ground-skid dry coolers — OFF-GRID 3× 750 kW NG gensets (N+1) · 1 MW / 2 MWh BESS · 800VDC-ready · Permian Basin, TX.
432 GPUs (6× NVL72) · ~124 TB HBM — a frontier open-weight AI factory in one box.
On July 27, 2026, Moonshot AI released the weights of Kimi K3 — a 2.8-trillion-parameter mixture-of-experts model, the largest open-weight release ever, benchmarked at the level of the most advanced American models (top-3 across the major intelligence indices).
And it did not stop: DeepSeek V4-Pro (MIT), GLM-5.3-Flash (MIT, 1M context), Qwen3.8-Max, Hunyuan Hy4 — downloadable frontier-class weights now land every week. A frontier model is no longer an API subscription. It is a file you download — and it serves from a single rack-scale footprint: a ~$0.5M-class hardware entry point (ESTIMATE), not a hyperscale build.
That changes who buys infrastructure. Every company can now run its own AI — weights on its own metal, inside its own perimeter. It doesn't watch you. It doesn't train on your data. It doesn't leak to competitors. And it answers to no one else's roadmap, rate limit, or deprecation notice.
But owning the model means owning the physics: 135–155 kW of heat per rack that never touches air, and a megawatt of power the grid interconnect queue won't deliver for years. That load doesn't go in a server closet — it goes in a purpose-built module. The 40 ft self-contained liquid-cooling container is the deployment unit of the open-weight era: set it on a pad, feed it pipeline gas, and a private frontier AI is on-prem — no interconnect queue, no stick-built construction.
Sources — Moonshot AI release materials · VentureBeat, 2026-07-16 · Artificial Analysis & Vals AI indices · model cards as of Aug 2026 · hardware entry point = ESTIMATE, pending config
Power is the bottleneck, not silicon: interconnect queues plus stick-built construction run to years. This module is factory-built to a ~90-day target and sited at the fuel — pipeline gas in, frontier tokens out.
| Model | Size (total / active) | Context | License | Role on this box |
|---|---|---|---|---|
| Kimi K3 | 2.8T / ~104B | 1M | custom K3 | open-weight intelligence ceiling — one prestige replica, not twenty |
| DeepSeek V4-Pro | 1.6T / 49B | 1M | MIT | most permissive giant — workhorse for private fine-tunes |
| Qwen3.8-Max | 2.4T / 95B | 262K–1M | custom | board-topping generalist — text weights public |
| Hunyuan Hy4 (preview) | 770B / 49B | 1M | open — confirm SPDX | fresh coding/research/finance — pin a canary replica |
| GLM-5.2 / 5.3 | ~750B / 40B | 1M | MIT | agent/coding family workhorse |
| GLM-5.3-Flash | 320B / 18B | 1M | MIT | default public endpoint — ~300 GB FP8, 50–100+ replicas per MW |
| DeepSeek V4-Flash | 284B / 13B | 1M | MIT | highest tokens/watt in the frontier set |
| Qwen3.8-27B | 27B dense | 262K | Apache 2.0 | routers, RAG, vision — fits 1–2 GPUs |
| gpt-oss-120b | 117B / 5.1B | — | Apache 2.0 | Western reasoning tier, easy compliance story |
A sane split on one module: 1–2 racks single-replica flagship (K3 / Qwen3.8-Max class) · 2–3 racks training + LoRA partition (V4-Pro, GLM-5.x, Hy4) · remaining power Flash autoscaled for API · a few nodes for 27B-class routers, embeds, classifiers · hold 10–15% headroom for KV.
License reality check: MIT / Apache tiers are safest to sell on. K3, Qwen3.8-Max, Llama 4 are open weights with commercial clauses — read the card before you put them behind a paid API.
Rev P2 engineering package for a containerized GPU data-center module. Full-liquid W45 direct-to-chip cooling — 100% liquid, fanless, no RDHx — is the default basis.
DI / 25% PG to the rack cold plates — 40 °C supply, 52–55 °C return. 1182 LPM on a DN125 header.
25% PG glycol — 45 °C supply, 57 °C return to the dry coolers.
The 40 ft High-Cube envelope is retained for the overhead tray/pipe zone. The end elevation shows the skid on grade beside the container with the FWS penetration.
Every dimension and process value lives in one Documentation file. The 3D model, 2D drawings, ISA P&ID and IFC all import it — nothing is hard-coded twice.
A Layout object computes all positions once, so model and drawings cannot drift. Zone boundaries, aisles and pad extents are computed in params, never drawn by hand. Change a value → re-run → the whole package updates consistently.
Consistency is structural, not manual: an impossible combination (e.g. BESS off + off-grid NG) is a hard error at generation time, not a shipped drawing.
Earlier “liquid-cooled” servers were hybrid: cold plates on GPU/CPU only. Memory, NICs, NVSwitch, PSUs and VRMs still rejected heat to air → fans, perforated bezels, RDHx.
The unsolved problem: cool 100% of the board on liquid. NVIDIA’s Rubin answer: liquid to every chip via a single tray inlet/outlet → sealed front, no fans, 6U→2U.
2026 convergence: the industry is trending 100% cold-plate fanless at 300 kW/rack, with the 100%-cold-plate class forecast at >10M units by 2027 — cold-plate beats immersion ≈ 95:5. Our baseline is exactly the architecture NVIDIA’s “Liquid Cooling AI Factories” describes as the previous, unsolved state.
Our original GB300 basis (~90% liquid + ~10% air → 6× RDHx) was that hybrid approach — at Rev P2 the 100%-liquid answer is the default.
| Metric | hybrid (legacy) | full_liquid — P2 default |
|---|---|---|
| Architecture | ~90% liquid + air | 100% liquid, fanless |
| RDHx / fans | 6 / required | 0 / none |
| Liquid load | 729 kW | 810 kW |
| TCS flow · header | 1062 LPM · DN125 (std) | 1182 LPM · DN125 |
| Total rejection | ~850 kW (heat balance) | ~850 kW (heat balance) |
| Front service aisle | 920 mm — FLAG | 1100 mm — OK |
Removing the 180 mm RDHx opens the aisle 920 → 1100 mm — every clearance clears.
Energy. ~4% cooling-cost cut per +1 °C loop temperature; >$4M/yr at 50 MW → order ~$65k/yr here (climate-dependent).
Waste heat. The 40–57 °C FWS loop is offtake-grade for district / greenhouse heat reuse (site-dependent).
Claim language. “Closed-loop with minimal adiabatic trim” — never “water-free.”
Part of the Rev P2 basis: a 1 MW / 2 MWh LFP skid off the electrical end, sited at 3.0 m from the compute container per NFPA 855 (2026). Shown in the 3D model, GA, IFC and isometric; storage node + ride-through note on the single-line.
Chemistry, cell selection and price remain ESTIMATE pending vendor quote.
| Entity | IFC class |
|---|---|
| IT racks | IfcBuildingElementProxy |
| Plate-HX (CDU integral) | IfcHeatExchanger |
| CDUs | IfcUnitaryEquipment |
| FWS pumps | IfcPump |
| Dry coolers | IfcCoolingTower |
| TCS / FWS headers | IfcPipeSegment |
| Electrical gear | IfcElectricDistributionBoard |
| NG gensets (×3) | IfcElectricGenerator |
| BESS | IfcElectricFlowStorageDevice |
Container interior = IfcSpace; floor/roof/walls = IfcSlab / IfcWall. Pset_ConceptBasis carries the Rev P2 basis: CDU CoolIT CHx1000, heat balance, future-proofing. Every shaped entity meshes with valid geometry.
The full-liquid default removes the 180 mm RDHx → front aisle 920 → 1100 mm. The legacy hybrid variant (./out_hybrid) retains its historical 920 mm / rear-aisle flags — one driver for the P2 default flip.
| Check | Target (mm) | Actual (mm) | Result |
|---|---|---|---|
| Electrical bay NEC 110.26 (480 V) | 1067 | 1500 | OK |
| IT front service aisle | 1067 | 1100 | OK |
| Overhead tray/pipe zone | 300 | 346 | OK |
| Rear access (no RDHx) | 0 | 50 | OK |
| Length far-end spare | ≥0 | 2111 | OK |
| NG exhaust → dry-cooler intake (ESTIMATE → CFD-screened) | 7600 | 7600 | OK |
The 7,600 mm exhaust→intake separation was a 25 ft rule-of-thumb. It is now screened by CFD: a steady RANS model (OpenFOAM 2512, simpleFoam, k–ε) of the whole pad — container, BESS, three coolers, three gensets — with a passive exhaust tracer released at each roof louvre.
Case + figures regenerate from cfd/gen_case.py → OpenFOAM (docker) → cfd/make_figures.py · exhaust louvre size, flow and temperature = ESTIMATE pending genset datasheet
The self-contained 40 ft module is the product — scale a site by replicating it.
The open ESTIMATE flags: dry-cooler count basis (3/module), CDU capacity-at-approach, BESS price, PUE band, NG-genset lead time / site derate / prime-duty certification.