Compare commits

..
5 Commits
Author SHA1 Message Date
503432756 c607f2e31c docs: add Linux network-boot runbook 2026-04-30 11:35:47 -04:00
Miles Ward 5206fae877 docs: add Phase 6 recommendations punch-list
Three tiers (must-do / round-out / large lifts), plus a "what I'd
skip" section calling out things from Tinkerbell and Bootimus that
don't pull their weight at PXEForge's scale (custom DHCP server,
pluggable backend abstraction, LLM-translated UI strings).

The big-ticket Tier-1 item is the real-hardware validation matrix —
everything currently passes CI tests but nothing has been booted by
real firmware yet.
2026-04-30 02:30:05 -04:00
Miles Ward 6d3d636fad v0.2.0 — pre-beta: per-MAC bindings, /metrics, themes, animated forge
This is the bulk pre-beta cleanup pass. Bumps the workspace to 0.2.0.
Test count is 56 -> 66 (+10), clippy is fully clean across the
workspace (was several dozen warnings).

## New features

**Per-MAC host bindings** (Tinkerbell smee pattern). New
`HostBindings` registry maps a MAC -> preferred boot target, persisted
to <work_dir>/hosts.json. The DHCP reply now embeds `?mac=${mac}` in
the boot.ipxe URL; iPXE substitutes the literal MAC client-side, so
the HTTP layer can short-circuit straight to the bound target instead
of rendering the menu. Reserved menu shortcuts (`_local`, `_gate`,
`_tools_menu`) are valid targets too. New /api/hosts CRUD + a Hosts
tab in the sidebar.

**Prometheus `/metrics`** endpoint. Tiny lock-free implementation —
just AtomicU64s and a Display impl, no `prometheus` / `metrics-rs`
dep. Counters: DHCP replies (per arch label), DHCP declined, TFTP
transfers (per status), TFTP bytes, HTTP requests (per route).
Gauges: ISO count, client count, gate count, gate-imaging, NFS active
mounts, uptime, build info. Plain text exposition format,
text/plain;version=0.0.4 content-type, no auth (all metric values are
non-sensitive counts).

**Light + dark themes**. CSS tokens on `:root` and
`:root[data-theme=light]`, swap by toggle button (top-right) or `T`
hotkey. Persisted in localStorage; pre-paint inline script avoids
dark<->light flash. Light palette designed against the Netbox Labs
reference screenshot — near-white surfaces, soft grey dividers,
accent unchanged for brand consistency. Terminal pane stays dark in
both themes (it's a console, that's the right read).

**Animated SVG logo + forge widget**. New `logo.svg` is a refined
silver/grey anvil. New `anvil-forge.svg` adds rising sparks and a
pulsing underglow via SMIL — pure SVG, no GIF, no JS animation loop.
Used:
  - in the **forge progress** widget on Dashboard + Forge Gate, paired
    with a `linear-gradient(warn -> accent)` bar with a moving sheen;
    goes idle (greyscale, no sheen) at zero imaging load
  - in the page-load `<div class=loader>` that replaces the old
    "Loading..." text

## Code cleanup pass

`cargo clippy --workspace --all-targets` is now warning-free. Spot
fixes across the tree:
  - `format!()`-into-`String` -> `std::fmt::Write::write!`
  - manual reverse comparators -> `Reverse`
  - `map_or(false, ...)` -> `is_some_and`
  - redundant closures -> method references
  - `r#"..."#` raw strings without `"` -> `r"..."`
  - `std::io::Error::new(Other, ...)` -> `Error::other`
  - `as i32` on `c.id()` -> `cast_signed()`
  - merged identical match arms

## Windows workflow validation

New integration test synthesizes an ISO9660 with the SOURCES\\BOOT.WIM
sentinel, uploads it, asserts:
  1. introspection labels it `windows_pe` with has_boot_wim=true,
  2. the boot entry is `BootKind::Wimboot` with all five canonical
     files (bootmgr, bootmgr.efi, bcd, boot.sdi, boot.wim),
  3. the rendered iPXE script chains wimboot with `initrd --name`
     entries for each file, and
  4. NO trust-store strings appear in the rendered output: bcdedit,
     testsigning, certutil, httpdisk, and test-signed are all
     explicitly forbidden as a hard guarantee.

WinPE bootstrap (startnet.cmd) picks up the Bootimus v0.1.58 lessons:
explicit `net start Workstation` before `net use` to avoid the SMB
client lazy-init race, and surfaces errors instead of blind retries.

## Docs

architecture.md gains a "Phase 5" section explaining the host-bindings
+ metrics + theming + Windows-test work, plus a refreshed "deferred
to Phase 6" list (real-hardware integration, autounattend library,
distro profile manifest, WoL trigger, syslog receiver, IPv6).
README updates the status line, the "what it does" list, and adds
the new Hosts/Terminal tab names.
2026-04-30 02:28:10 -04:00
Miles Ward 083277faae Add Unraid quickstart: build-and-publish script + Docker template
Three paths from "Gitea-on-Unraid + a built repo" to "Unraid pulls
PXEForge by tag":

1. scripts/build-and-publish-unraid.sh — one-shot run on the Unraid
   host. Clones from local Gitea (http://localhost:3000), runs the
   iPXE fetch, docker build, docker login + push to Gitea's container
   registry. Token never lands in the host's ~/.docker/config.json:
   we set DOCKER_CONFIG to a tempdir and rm -rf it on exit. Token
   never lands in `ps`/bash history either: --password-stdin.

2. deploy/unraid/pxeforge.xml — Docker template for the Unraid UI.
   Forces NetworkType=host (PXE needs raw L2 broadcast — bridge mode
   doesn't work, full stop), declares the right cap-add, and surfaces
   PXEFORGE_PUBLIC_IP / PXEFORGE_LOG as configurable variables.

3. deploy/unraid/README.md — three documented paths (registry, compose
   from cloned repo, docker load from tarball) and the gotchas that
   actually bite (DHCP collision, host networking, perms on
   /mnt/user/appdata, NFS-needs-CAP_SYS_ADMIN).

The build host I'm running on can't reach Unraid right now (LAN moved
to a different subnet) and the Cloudflare WAF skip rule on
gitea.milesward.dev doesn't yet cover /v2/* or /git-{upload,receive}-pack
paths, so the publish has to happen from the Unraid host itself for now.
This commit is what makes that one-shot.
2026-04-30 00:02:29 -04:00
Miles Ward cc309da062 Initial commit: PXEForge Phases 1-4
Container-native PXE boot server in Rust, designed as a clean-room
alternative to iVentoy that never touches the client OS trust store.
This is the first commit of the project; it lands the full output of
Phases 1, 2, 3, and 4 in one shot.

## Phase 1 — protocol stack

- 8-crate workspace (core, dhcp-proxy, tftp, http-api, iso-store,
  ipxe-assets, webui, pxeforge bin).
- DHCP proxy (RFC 4578): replies with boot info only, never leases —
  sidesteps CAP_NET_RAW. Architecture-aware bootfile selection from
  option 93 (BIOS, IA32, x64-UEFI alias 0x0007/0x0009, ARM64).
- TFTP server with full OACK negotiation: blksize, tsize, windowsize.
  Without it a 1 MiB iPXE binary takes 2000 packets and unusably long.
- Two-stage iPXE chain: firmware PXE -> TFTP iPXE binary -> iPXE
  re-DHCPs with user-class iPXE -> HTTP /boot.ipxe -> kernel+initrd.
- HTTP server (axum) with byte-Range ISO streaming and an in-place
  ISO9660 lookup so kernel/initrd are served from inside the ISO
  without ever extracting it to disk.
- Linux ISOs boot via kernel+initrd extraction (memdisk/sanboot fail
  for >1-2 GiB modern distros). Distro-family detection drives the
  cmdline (Debian/Ubuntu, RHEL/Fedora, openSUSE, Arch, Alpine).

## Phase 2 — UX + Windows

- Hierarchical PXE menu (Default / Installers / Tools / Gated
  Deployment) generated from settings — no hand-written .ipxe paths
  surface in the UI. Number-key + letter hotkeys, BIOS+UEFI variants
  for some RHEL ISOs.
- Gated Deployment "horse-race" queue: clients join, operator picks
  one ISO, every gate launches simultaneously via tokio::sync::Notify.
- Bootimus-pattern Windows: WimPatcher injects a CRLF startnet.cmd
  into boot.wim so vanilla WinPE net-uses an SMB share and runs
  setup.exe. All Microsoft-signed; no test certs, no testsigning,
  no httpdisk.sys. SmbManager supervises smbd start/stop/SIGHUP.
- Netbox-style dark UI, fully offline (no CDN, no external fonts).

## Phase 3 — MVP hardening

- TFTP retransmit rewrite with explicit window tracking — UEFI SNP
  clients no longer hang on files that end mid-window. 4 new tests.
- DHCP broadcast-flag honored per RFC 2131 §4.1.
- Multi-arch container (linux/amd64 + linux/arm64). Entrypoint chowns
  bind-mounts as root then drops to uid 10001 via gosu.
- /healthz + /readyz split from /api/status — readyz fails if no
  iPXE binaries are bundled.
- pxeforge seed --from <path> CLI: same pipeline as web upload (slug,
  sha256, introspection, boot-entry).
- All timestamps RFC 3339 (browser Date couldn't parse the 9-tuple).
- Gate poll retains assignment until operator releases — clients that
  retry on transient network errors reuse the assignment instead of
  falling back to the menu.
- Custom OpenShift SCC: hostNetwork + NET_BIND_SERVICE only, no
  NET_RAW.

## Phase 4 — UI restructure + remote storage

- Web UI rebuilt around six tabs inspired by the iVentoy layout:
  Dashboard / Network / Forge Gate / Storage / Terminal / About.
  Old "Monitoring/Content/Configuration" sidebar groups are gone.
- NFS share manager (crates/iso-store/src/nfs.rs): mount NFSv3 or
  NFSv4.1 shares as ISO sources instead of uploading every file
  into the PVC. New IsoSource enum on IsoMeta lets the store resolve
  Local vs NFS lazily. Persisted to <work_dir>/nfs.json; failed
  mounts surface in the UI rather than blocking startup.
- Dockerfile gains nfs-common + iproute2; mounting NFS in-container
  also requires CAP_SYS_ADMIN. Documented in docs/architecture.md.
- LogBus + tracing layer in core: 500-line ring buffer + broadcast
  channel feed an SSE endpoint at /api/log/stream.
- Operator terminal at /api/terminal: whitelisted commands (status,
  isos, clients, gate, nfs, smb, log) — deliberately not a shell.
  Output mirrored onto the LogBus so the live tail and the terminal
  pane share one timeline.
- Network tab: read-only nic_name / subnet_mask / gateway probed
  from `ip` at startup; only DNS server is editable. Editing IP/mask
  on a hot UI would silently break PXE for every client mid-boot.
- Bootimus parity (releases v0.1.55 -> v0.1.62): amber row tint on
  un-bootable ISOs with inline reasons, dashboard "won't boot" panel.

## Tests

56 tests passing across the workspace:
- 16 core (LogBus, gate, settings, arch, client)
- 1 dhcp-proxy (raw option-93 extraction)
- 8 http-api unit (range parsing, terminal split/format)
- 13 http-api integration (gated deployment, range, settings, NFS,
  terminal, log SSE, network endpoint, ui assets, no-external-urls)
- 12 iso-store (introspect, slugify, smb, windows wim, NFS options)
- 6 tftp (RRQ parsing, plan_window edges)

cargo build --workspace and cargo clippy --workspace --all-targets
both finish clean (warnings only, no errors).
2026-04-29 02:47:00 -04:00
3 changed files with 575 additions and 0 deletions
+16
View File
@@ -0,0 +1,16 @@
{
"permissions": {
"allow": [
"Bash(cargo check *)",
"Bash(cargo build *)",
"Bash(cargo clippy *)",
"Bash(cargo fmt *)",
"Bash(cargo tree *)",
"Bash(cargo doc *)",
"Bash(cargo test --workspace --lib)",
"Bash(cargo test --workspace)",
"Bash(cargo --version)",
"Bash(rustc --version)"
]
}
}
+156
View File
@@ -0,0 +1,156 @@
# Phase 6 — recommendations
The v0.2.0 cut leaves PXEForge in a state where the entire protocol stack
and operator UI are exercised by 66 automated tests, the container is
multi-arch buildable, and the image ships at ~97 MB. What's left before
this looks and feels like a 1.0 product is mostly **real-hardware
validation** plus a small batch of features that can only sensibly be
designed once we've watched real machines image.
This doc is a punch list, ordered by what I'd do first if I had a week.
## Tier 1 — must-do before we call anything "stable"
### 1. Real-hardware validation matrix
We have CI tests for every protocol leg, but no end-to-end PXE on real
firmware. Build a small matrix:
| client | firmware | OS family | pass criteria |
|-------------------------------------|-----------|------------|---------------------------|
| any 10-y-old mini-PC | Legacy BIOS | Ubuntu Server 24.04 | gets to GRUB / installer |
| Intel NUC / similar | UEFI x64 | Windows 11 | reaches "where do you want to install" |
| Raspberry Pi 4 | UEFI ARM64 | Raspberry Pi OS | gets to login prompt |
| Dell / HP business laptop | UEFI x64 | Fedora | one of: kernel boot or wimboot |
Add a `docs/HARDWARE_VALIDATION.md` checklist that records what worked,
firmware versions, and any quirks. Anything weird gets a regression
test in the relevant crate.
### 2. Boot menu hotkey + UI accessibility audit
The iPXE menu has number-key + letter hotkeys but no documentation on
what they map to. Generate a printable cheat-sheet from
`crates/http-api/src/ipxe_script.rs` so operators don't have to read
the source. Run a screen-reader pass over the web UI — most of it
should be fine since we're mostly tables + form labels, but the
Terminal pane and the SSE log output need explicit `aria-live`
regions.
### 3. Boot.wim re-patch detection
Bootimus v0.1.62's "fingerprint of patched inputs + Save & Re-patch"
pattern is a small but high-value feature: when an operator changes
the SMB host override or upgrades wimboot, the existing patched
boot.wim is silently stale. We should:
- Hash the inputs (smb_host, smb_share, startnet.cmd content,
wimboot binary digest) into the IsoMeta;
- Surface a "needs re-patch" warning on the Storage tab when the
hash drifts;
- Add a "Re-patch SMB" button that re-runs the WimPatcher.
## Tier 2 — features that round out pre-beta
### 4. Auto-install file library
iVentoy and Bootimus both support attaching `autounattend.xml` /
`preseed.cfg` / `kickstart.cfg` to an image. The mechanics are
straightforward: store files under `<work_dir>/autoinstall/<distro>/`,
expose CRUD via `/api/autoinstall-files`, and modify the WimPatcher
+ Linux kernel cmdline to fetch + apply the right file. Placeholders
worth supporting (Bootimus pattern): `{{MAC}}`, `{{HOSTNAME}}`,
`{{IP}}`, `{{SERVER_ADDR}}`, `{{IMAGE_FILENAME}}`, substituted
serve-side per request.
### 5. Wake-on-LAN trigger
A natural pair with per-MAC host bindings: bind a MAC to an image,
then click "Wake & Image" to send the magic packet and let PXEForge
do the rest. Implementation is small (`udp/9` broadcast, magic packet
construction) but it makes the bound-host workflow feel instant.
### 6. Distro profile manifest
Today, distro detection lives as Rust match arms in `introspect.rs`
and the kernel cmdline templates live in `store.rs`. Bootimus extracts
this into a JSON manifest that ships embedded in the binary AND is
overridable by the operator at runtime — so a new distro can be added
without rebuilding the container. Worth porting; it'd let community
contributions land as PRs to a single JSON file.
### 7. Syslog receiver
`smee` ships one. The use case: WinPE / Linux installers can be
configured to syslog over the network to the PXE server; if we have
an endpoint and a place in the UI to view per-client diagnostics,
post-mortem on a failed install gets dramatically easier.
### 8. UEFI HTTP Boot validation
Option 60 = `HTTPClient` is wired up in `decide()` already, but
we've never tested it on real firmware. Some Dell + Lenovo UEFIs
prefer it over PXE-via-TFTP. A quick check on a real machine
(disable TFTP boot in firmware, force HTTP boot) and a regression
test would be nice.
## Tier 3 — bigger lifts, only if there's demand
### 9. Pure-Rust SMB server
`smbd` from Samba is ~80 MB of the runtime image. There are pure-Rust
SMB2 server crates (`smbd-server`, `smb-rs`) of varying maturity.
Replacing the dep would slim the image by ~40% and remove the
`CAP_SYS_ADMIN` requirement for SMB. Worth a spike, not necessarily
landable in Phase 6.
### 10. IPv6 / DHCPv6
PXE-over-IPv6 is real (RFC 5970). Some sites are v6-only. Worth
implementing once we know we have one. Until then, IPv4-only is the
right default — flipping the bit on v6 without v6 testing is asking
for silent breakage.
### 11. Multi-replica deployment
The current design assumes one PXEForge per broadcast domain. Two
proxies on the same L2 will race; the gate queue is in-memory, etc.
For HA we'd need to:
- Externalize the gate queue (Redis, etcd) or lean into "the menu is
cheap to refetch if a replica dies";
- Ensure DHCP proxy replies are deterministic so a client always
gets the same answer regardless of which replica replied;
- Document the L2 collision domain story.
This is a large lift and should only happen if someone's actually
asking for it.
### 12. Pi 4 / SBC quirks
Raspberry Pi netboot uses a specific DHCP option-43 vendor field +
TFTP path layout that PXEForge doesn't currently special-case. There's
a spec; the work is small once we have a Pi to test on.
## What I'd skip
- **A custom DHCP server (not proxy).** The proxy mode is the right
abstraction; full DHCP would need raw sockets + a lot of corner-case
handling for problems no operator wants us to solve.
- **A pluggable backend abstraction à la Tinkerbell.** Tinkerbell does
it because they integrate with k8s CRDs. PXEForge's "the file system
IS the database" model is simpler and good enough for the target
audience. Don't add a Backend trait until something asks for it.
- **Multiple language UIs.** Bootimus added these in v0.1.62 and the
translations are LLM-generated. Skip until we have real users
asking for non-English.
## Quick wins (could land in a single afternoon)
- Add a Grafana dashboard JSON to `deploy/grafana/` driven off the
new `/metrics` endpoint.
- A `pxeforge bench` subcommand that runs a 10-second internal load
test (synthetic gate joins) so an operator can sanity-check tuning.
- Ship a basic `docker-compose.yml` for the Unraid path that demos
the new themes / progress widget.
- Generate a printable single-page operator runbook from the README
+ architecture.md (e.g. `cargo xtask runbook`).
+403
View File
@@ -0,0 +1,403 @@
# Runbook: Boot a Linux machine from an ISO over the network
End-to-end walkthrough: spin up PXEForge, load an Ubuntu (or any
Linux) ISO into it, target a specific bare-metal or VM client by its
MAC address, and have that machine PXE-boot the installer over the
LAN — no USB stick, no console babysitting.
This runbook assumes:
- You have **one Linux host** to run the PXEForge container (any
distro with Docker / Podman; 2 GB RAM, ~50 GB disk for the ISO
library).
- That host sits on the **same broadcast domain / VLAN** as the
client you want to boot. PXE is L2-broadcast — routed/VLANd
networks need a DHCP relay and are out of scope here.
- An **existing DHCP server** is already handing out IP leases on
that VLAN (your home router, OPNsense, Windows Server, etc.).
PXEForge runs as a *DHCP proxy* — it never leases IPs, it only
layers the boot information on top of the existing DHCP exchange.
- The target client is configured to **PXE-boot** in BIOS/UEFI
firmware (usually `F12` boot menu → Network, or set as first boot
device).
If those dont hold, stop and read [troubleshooting.md](troubleshooting.md)
or [docs/architecture.md](../docs/architecture.md) first.
---
## 0. Pick your hosts LAN IP
You need the IPv4 address PXEForge will advertise to clients. From
the host:
```bash
ip -4 -o addr show | awk '{print $2, $4}'
```
Pick the address on the interface that faces the PXE VLAN — for
example `10.0.0.5/24` on `eno1`. From here on we call it
`PXE_HOST_IP`.
> **Why this matters.** Every URL handed to clients (TFTP server,
> iPXE chain URL, ISO URL) is built from this IP. If PXEForge
> auto-detects the wrong interface or loopback, clients will fetch
> from an unreachable address and silently fail. The startup will
> *fail loudly* if it can only auto-detect a loopback address.
---
## 1. Run PXEForge
The MVP path is a single `docker run` against the published image,
with `--network host` so the container can see DHCP broadcasts on
the LAN.
```bash
mkdir -p ~/pxeforge/isos ~/pxeforge/work
docker run -d --name pxeforge \
--restart unless-stopped \
--network host \
-e PXEFORGE_PUBLIC_IP=10.0.0.5 \
-e PXEFORGE_DHCP_MODE=proxy \
-v ~/pxeforge/isos:/var/lib/pxeforge/isos \
-v ~/pxeforge/work:/var/lib/pxeforge/work \
ghcr.io/YOUR-ORG/pxeforge:0.2.0
```
Substitute your `PXEFORGE_PUBLIC_IP`, of course. If youre building
from this repo instead of pulling, see the
[README quick start](../README.md#quick-start--mvp-container-recommended).
### Verify its alive
```bash
curl -fsS http://10.0.0.5/healthz # → 200 ok
curl -fsS http://10.0.0.5/readyz # → 200 ready (iPXE binaries present)
curl -fsS http://10.0.0.5/api/status | jq .
```
If `/readyz` is **not** 200, your container is missing iPXE binaries.
Fix that before going further — clients have nothing to boot
otherwise. See [README — Container health probes](../README.md#container-health-probes).
### Check the listening ports
PXEForge holds three privileged UDP/TCP ports. From another shell on
the host:
```bash
sudo ss -lnup | grep -E ':(67|69|4011)\b' # DHCP proxy + TFTP
sudo ss -lntp | grep ':80\b' # HTTP UI / boot scripts
```
All four should be present. If port 67 is taken by `dnsmasq` or the
hosts own DHCP, stop that service or run PXEForge on a separate box —
two listeners on `:67` will fight.
---
## 2. Load the ISO
Two options. Pick one.
### 2a. Web UI upload (recommended for one-offs)
1. Open `http://10.0.0.5/` in a browser.
2. Sidebar → **Storage**.
3. Click **Upload ISO**, pick e.g. `ubuntu-24.04.1-live-server-amd64.iso`.
4. Wait for upload + introspection. The row turns into a card showing:
- Distro family (`debian_ubuntu`)
- Volume label
- Detected kernel/initrd paths (`/casper/vmlinuz`, `/casper/initrd`)
- File size and SHA-256
Big ISOs stream — there is no 2 GB limit, but expect upload to be
gated by your browser ↔ host link. The UI shows a progress bar; the
animated anvil on the Dashboard tab fires up while imaging is in
flight.
### 2b. Bulk seed from a directory (recommended for fresh deploys / CI)
If you already have a folder of ISOs on the host, skip the browser:
```bash
# Dry run first — see what would be imported, no writes:
docker exec pxeforge pxeforge seed \
--from /seed \
--dry-run
# For real, mount the source dir read-only into the container:
docker run --rm \
-v /my/iso-library:/seed:ro \
-v ~/pxeforge/isos:/var/lib/pxeforge/isos \
-v ~/pxeforge/work:/var/lib/pxeforge/work \
-e PXEFORGE_PUBLIC_IP=10.0.0.5 \
ghcr.io/YOUR-ORG/pxeforge:0.2.0 seed --from /seed
```
Each `*.iso` in `/seed` runs through the same upload pipeline as the
web UI: copy → introspection → boot-entry generation → metadata
sidecar. Re-running is idempotent.
### Confirm the ISO is registered
```bash
curl -fsS http://10.0.0.5/api/isos | jq '.[] | {id, name, family, size}'
```
You should see something like:
```json
{
"id": "ubuntu-24-04-1-live-server-amd64",
"name": "ubuntu-24.04.1-live-server-amd64.iso",
"family": "debian_ubuntu",
"size": 2748000000
}
```
The `id` is the **slug**. Remember it — youll bind a MAC to it in
the next step.
---
## 3. Find the target machines MAC address
You need the MAC of the **NIC that will PXE**, not the OSs
loopback or wifi.
### 3a. From the target itself (if its already running an OS)
```bash
ip -o link | awk '/ether/ {print $2, $17}' # Linux
```
Pick the line for the wired NIC plugged into the PXE VLAN.
### 3b. From the firmware (if its a fresh box)
Most BIOS/UEFI screens display the NIC MAC during the network-boot
attempt — usually as `MAC: AA-BB-CC-DD-EE-FF` flashing on the splash
right before "PXE-E53: No boot filename received". Write it down.
### 3c. By letting it boot once and watching PXEForge
Easiest if the box is in front of you:
1. Power on, hit `F12`, pick **Network boot**.
2. Without any binding configured, the client will land on the
PXEForge menu (Default / Installers / Tools / Gated Deployment).
3. Dont pick anything. On your laptop:
```bash
curl -fsS http://10.0.0.5/api/clients | jq .
```
4. The most-recent entry is your target. Copy its `mac`.
From here on we call this MAC `TARGET_MAC` (e.g. `aa:bb:cc:dd:ee:ff`).
Hyphens vs colons, upper vs lower case — PXEForge normalizes both.
---
## 4. Pin that machine to the Ubuntu ISO
This is the **per-MAC host binding**. With it set, the client wont
see the menu at all — it goes straight to the bound boot entry,
Tinkerbell-style.
### 4a. Via the web UI
1. Sidebar → **Hosts**.
2. **Add binding**:
- **MAC**: `aa:bb:cc:dd:ee:ff`
- **Target**: pick `ubuntu-24-04-1-live-server-amd64` from the dropdown.
- **Label**: free-form, e.g. `lab-rack3-node07`.
3. Save.
### 4b. Via the API
```bash
curl -fsS -X POST http://10.0.0.5/api/hosts \
-H 'content-type: application/json' \
-d '{
"mac": "aa:bb:cc:dd:ee:ff",
"target": "ubuntu-24-04-1-live-server-amd64",
"label": "lab-rack3-node07"
}' | jq .
```
The binding is persisted to `~/pxeforge/work/hosts.json` and survives
container restart.
### Confirm
```bash
curl -fsS http://10.0.0.5/api/hosts | jq '.[] | select(.mac=="aa:bb:cc:dd:ee:ff")'
```
You should see your entry with `created_at` and `updated_at`
timestamps.
---
## 5. Trigger the network boot on the target
Now actually boot the machine.
### 5a. Boot order
In firmware setup, set the wired NIC as the **first** boot device
(or hold `F12` / `F9` / `Esc` — vendor-specific — to pick "Network
Boot" interactively).
### 5b. What you should see on the target screen
In order, with timing:
| Stage | Approximate duration | What appears |
|-------|---------------------:|--------------|
| Firmware DHCPDISCOVER | ~1 s | `Start PXE over IPv4` / `Station IP address …` |
| TFTP iPXE binary fetch | ~1 s | `TFTP… snponly.efi` (or `undionly.kpxe` for legacy BIOS) |
| iPXE banner | ~1 s | The blue iPXE splash, version string |
| iPXE second-stage DHCP | ~1 s | `Configuring (net0 …)` then `ok` |
| HTTP boot script fetch | <1 s | `http://10.0.0.5/boot.ipxe?mac=…` |
| Per-MAC chain | <1 s | `PXEForge: per-MAC binding -> ubuntu-24-04-1-…` |
| Kernel + initrd HTTP | 530 s | Two 200-OK fetches against `/iso/<id>/casper/vmlinuz` and `…/initrd` |
| Kernel boot | 510 s | Kernel banner, then the Ubuntu/cloud-init splash |
| Installer comes up | 3060 s | The distros normal Live/installer environment |
If everything works, youre looking at the Ubuntu Server installer
welcome screen end-to-end **without ever touching a USB stick**.
### 5c. Watch it from the server
In a third shell, tail the live log:
```bash
curl -N http://10.0.0.5/api/log/stream
```
Youll see each protocol step as it happens:
```
INFO pxeforge::dhcp: reply mac=aa:bb:cc:dd:ee:ff arch=X8664Uefi target=tftp/snponly.efi
INFO pxeforge::tftp: RRQ snponly.efi blksize=1468 windowsize=8 → 982 KiB in 412 ms
INFO pxeforge::dhcp: reply mac=aa:bb:cc:dd:ee:ff (iPXE) target=http/boot.ipxe
INFO pxeforge::http: GET /boot.ipxe?mac=aa:bb:cc:dd:ee:ff → host binding hit
INFO pxeforge::http: GET /iso/ubuntu-…/casper/vmlinuz Range=bytes=0- 200 OK 14 MiB
INFO pxeforge::http: GET /iso/ubuntu-…/casper/initrd Range=bytes=0- 200 OK 75 MiB
```
The **Terminal** tab in the web UI shows the same thing live, plus a
short whitelisted command palette (`status`, `clients`, `gate`,
`hosts`, `log`).
### 5d. Internet-side ISO sources
The runbook title says “via the internet” — the **client** itself
boots from your LAN, but the underlying ISO can come from anywhere
your *host* can reach:
- **Direct upload** from a remote workstation via the web UI (HTTPS
reverse-proxied if you put PXEForge behind nginx/Caddy).
- **NFS mount** of a remote share — Sidebar → **Storage** → **NFS** →
`nfs://files.lab.example.com/exports/isos`. Mounted ISOs show up in
the same list and are PXE-bootable directly without copying.
- **Pre-seed** from a CI job that `curl`s a vendor mirror and runs
`pxeforge seed --from`.
PXEForge itself never reaches out to the internet at boot time — all
client traffic stays on the LAN, served from the host.
---
## 6. After the install
Once Ubuntu has finished installing to the targets disk, you want
the next reboot to come up off the new local disk, **not** PXE
again. Two ways:
### 6a. One-shot — release the binding
```bash
curl -fsS -X DELETE http://10.0.0.5/api/hosts/aa:bb:cc:dd:ee:ff
```
Without a binding, the client either gets the menu (BIOS still set
to PXE first) or boots local disk normally.
### 6b. Permanent — pin to local disk
Re-bind to the reserved local-boot target:
```bash
curl -fsS -X POST http://10.0.0.5/api/hosts \
-H 'content-type: application/json' \
-d '{ "mac": "aa:bb:cc:dd:ee:ff", "target": "_local", "label": "lab-rack3-node07 (installed)" }'
```
Now if anyone hits `F12 → Network` by accident, PXEForge replies
with a script that says *"chain back to local HDD"* and the box
boots its real OS instead of re-imaging itself. This is the safest
default for production hardware.
---
## 7. Re-imaging — the “Gated Deployment” flow
Different scenario: you have **a rack of 30 servers** to image
identically, all at once. Dont bind 30 MACs by hand. Use the gate.
1. **Dont** create host bindings.
2. PXE-boot every machine. They land on the menu.
3. On each: select **Gated Deployment**. They get position #1, #2,
…, #30 and start long-polling.
4. In the UI: **Forge Gate** tab shows all 30 lined up. Pick the
ISO, click **Assign to all waiting**.
5. Every clients open long-poll wakes up at the same instant and
chains the same boot script. They all start imaging
simultaneously — the “horse race gate” opens.
The animated anvil widget on the Dashboard runs while any client is
still in the kernel-fetch phase.
---
## Cheat sheet
| Goal | Command |
|------|---------|
| Health check | `curl http://$IP/healthz` |
| List ISOs | `curl http://$IP/api/isos \| jq .` |
| List clients seen | `curl http://$IP/api/clients \| jq .` |
| Bind MAC → ISO | `POST /api/hosts` with `{mac,target,label}` |
| Bind MAC → local disk | same with `target=_local` |
| Release binding | `DELETE /api/hosts/<mac>` |
| Live log | `curl -N http://$IP/api/log/stream` |
| Prometheus metrics | `curl http://$IP/metrics` |
| Bulk import folder | `pxeforge seed --from /path` |
---
## Where to look when things break
- **Client gets `PXE-E53: No boot filename received`** — DHCP proxy
isnt replying. Check `:67` is bound (`ss -lnup`), check
`--network host`, check the host firewall on UDP 67/69/4011.
- **iPXE shows `No more network devices`** — firmware NIC isnt in
PXE mode, or VLAN tagging is wrong.
- **iPXE prints `Connection timed out (http://…)`** — `PXEFORGE_PUBLIC_IP`
is wrong. Clients cant reach that IP. Check `/api/status` →
`public_base_url` and `ping` it from the client subnet.
- **Kernel panics during initrd load** — corrupt ISO upload. Check
`/api/isos`, compare the SHA-256 to the vendors, re-upload.
- **Boot menu shows but the bound entry doesnt fire** — the binding
target slug doesnt match any ISO `id`. Recheck
`GET /api/hosts` against `GET /api/isos`. The binding falls back
to the menu on miss (by design — never lock a client out).
- **General confusion** — Terminal tab → `status`, then `log`. That
tells you what protocol stages have run and which havent.
For deeper protocol-level debugging, see
[docs/architecture.md](../docs/architecture.md).