v0.7.4: probe-based introspection — remote shares classify, gparted bug fixed, Storage pagination

Introspection (the headline): detection is now probe-based. Instead of
grepping raw sectors for filename strings, we walk the ISO9660
directory tree and check whether the well-known boot files actually
exist — and the same probes run over NFS READ3 / SFTP seek-reads, so
share-hosted ISOs finally classify instead of registering as Unknown.

iso-store:
- New iso_fs module: the read-only ISO9660 walker (generalized from
  http-api) over an IsoReadAt trait — local files, NFS, SFTP, and the
  in-memory test images all share it. Iterative walk, 4 MiB directory
  cap, strict-mastering trailing-dot normalization (VMLINUZ.;1 now
  matches /vmlinuz), CachingReadAt collapses repeated directory reads
  during the probe pass (~60 → ~6 round-trips per remote ISO).
- introspect.rs rewritten (INTROSPECT_REV 2): PVD label → El Torito →
  /sources/boot.wim probe → verified Linux kernel+initrd probe table →
  local-only 16 MiB UDF-Windows scan → filename-token fallback.
  * Fixes the false-Windows bug: any Linux ISO shipping GRUB/syslinux
    chainload modules contains the literal "bootmgr", so gparted-live
    classified as WindowsPe. Linux probes now run first; the byte scan
    only sees ISOs nothing else claimed. Local ISOs re-probe once on
    startup via the rev bump — no re-upload.
  * Kernel entries are emitted only when kernel+initrd verifiably
    exist (no more guessed paths that 404 at boot). Debian-live /
    d-i netinst / CoreOS shapes classify for the UI but keep their
    working sanboot entries (their boot protocols need args we don't
    render yet; CoreOS additionally needs its embedded ignition).
  * Label + filename vocab extended: rhcos/coreos/openshift/okd,
    gparted/clonezilla/kali/tails, almalinux/rocky, sles, manjaro.
- NFS + SFTP managers: per-ISO IsoReadAt readers (READ3-at-offset with
  short-read looping / seek+read_exact), background introspection pass
  after each scan — entries register instantly with a provisional
  filename-based report (rev 0, optimistic sanboot preserved) and
  upgrade in place as probes land (30s/ISO timeout, failures keep the
  provisional). locate_in_iso() exposes the walker to the HTTP layer.
- remote_cache: introspection results persisted per protocol keyed
  share/path@size and gated on INTROSPECT_REV — container restarts
  re-probe only new/replaced ISOs; upgrades re-probe exactly once.
- SMB: smbclient can't seek, so SMB ISOs get the filename-token family
  (rev stays 0 → sanboot entry + "awaiting introspection" label).
- IsoStore::update_external_introspection swaps in completed reports
  and regenerates boot entries, preserving category/password.

http-api:
- /iso/{id}/{*path} now serves files from inside NFS/SFTP-hosted ISOs
  (remote ISO9660 lookup + ranged share stream) — verified kernel
  entries on remote Linux ISOs are actually bootable, end to end.
- iso_fs.rs deleted in favor of the shared iso-store module.
- full_flow fixtures build real directory trees via the shared
  test-image builder (new iso-store feature) — a label-only blob no
  longer earns a kernel entry, by design.

webui:
- Available images: paged 5 per page with a quiet footer pager
  (Showing X–Y of N · Prev/Next), filter-then-paginate, page resets on
  search input. Fifty images is five clean pages, not a scroll wall.
- Hosts/Queue profile: "Unattended file (in Storage → Advanced)" so
  the picker says where the files live.
- Row badge keys on introspect_rev: probed remote ISOs read like local
  ones; un-probed say "awaiting introspection".

Validation: clippy pedantic clean, fmt clean, 316 workspace tests
green (+17: walker, probe shapes incl. gparted regression + CoreOS,
filename table, cache round-trips), webui syntax-checked.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
This commit is contained in:
Miles Ward
2026-06-12 15:21:14 -04:00
co-authored by Claude Opus 4.8
parent 934cfbab46
commit 6524aa4118
18 changed files with 1810 additions and 401 deletions
+204 -5
View File
@@ -57,7 +57,9 @@
//! UI to ask for. (If a future server needs Kerberos or non-default
//! uid mapping we can add those, but for ISO read access nobody does.)
use crate::introspect::IntrospectionReport;
use crate::introspect::{introspect_reader, provisional_report};
use crate::iso_fs::{self, FileLocation, IsoReadAt};
use crate::remote_cache::RemoteIntrospectCache;
use crate::store::{generate_boot_entries_for, slugify_str, IsoSource, IsoStore};
use bytes::Bytes;
use nfs3_client::tokio::TokioConnector;
@@ -100,6 +102,18 @@ const READ_CHUNK_BYTES: u32 = 64 * 1024;
/// client park gigabytes of decoded ISO in RAM.
const STREAM_BUFFER_DEPTH: usize = 16;
/// v0.7.4: per-ISO budget for a background introspection probe. A probe
/// is one connection plus a few dozen KiB-sized reads — sub-second on a
/// LAN — so anything past this is a wedged server, not a slow one.
const INTROSPECT_TIMEOUT: Duration = Duration::from_secs(30);
/// One queued background-introspection unit (v0.7.4).
struct ProbeJob {
iso_id: String,
filename: String,
size: u64,
}
/// One configured NFS share. The id is derived from server+export so
/// re-adding the same coordinates is idempotent.
#[derive(Debug, Clone, Serialize, Deserialize)]
@@ -177,6 +191,9 @@ pub struct NfsShareManager {
/// opens its own NFS connection so concurrency isn't a hard
/// requirement, but serializing keeps log output predictable.
op_lock: Arc<tokio::sync::Mutex<()>>,
/// v0.7.4: persisted introspection results keyed `share/path@size`,
/// so a restart re-probes only new or replaced ISOs.
introspect_cache: RemoteIntrospectCache,
}
impl NfsShareManager {
@@ -190,6 +207,7 @@ impl NfsShareManager {
inner: Arc::new(Mutex::new(Inner::default())),
iso_store,
op_lock: Arc::new(tokio::sync::Mutex::new(())),
introspect_cache: RemoteIntrospectCache::open(work_dir, "nfs_introspect_cache.json"),
}
}
@@ -387,12 +405,26 @@ impl NfsShareManager {
};
let mut count = 0u32;
let mut to_probe: Vec<ProbeJob> = Vec::new();
for entry in listing {
let iso_id = format!("nfs-{}-{}", share.id, slugify_str(&entry.filename));
// Same approach as SMB: no real introspection over the
// network in v0.4.67. The boot-entry generator falls back
// to filename-based sanboot detection.
let report = IntrospectionReport::default();
// v0.7.4: real introspection over the share — NFSv3 READ3
// takes an offset, so the ISO9660 probes work remotely. A
// cache hit registers the full report immediately; a miss
// registers a provisional filename-based report (so the scan
// returns fast) and queues a background probe that upgrades
// the entry in place.
let cached = self
.introspect_cache
.get(&share.id, &entry.filename, entry.size);
let report = cached.unwrap_or_else(|| {
to_probe.push(ProbeJob {
iso_id: iso_id.clone(),
filename: entry.filename.clone(),
size: entry.size,
});
provisional_report(&entry.filename)
});
let boot_entries = generate_boot_entries_for(&iso_id, &entry.filename, &report);
let source = IsoSource::Nfs {
share_id: share.id.clone(),
@@ -413,11 +445,88 @@ impl NfsShareManager {
target: "openpxe::nfs",
id = %id, server = %share.server, export = %share.export,
iso_count = count,
pending_introspection = to_probe.len(),
"NFS share scanned"
);
if !to_probe.is_empty() {
self.spawn_introspection_pass(&share, to_probe);
}
Ok(count)
}
/// v0.7.4: probe each queued ISO over its own NFS connection and swap
/// the full introspection into the store as results land. Runs
/// detached so neither startup nor the share-add API call waits on a
/// 40-ISO library; per-ISO failures (or a share removed mid-pass)
/// leave the provisional entry in place, which still sanboots.
fn spawn_introspection_pass(&self, share: &NfsShare, work: Vec<ProbeJob>) {
let store = self.iso_store.clone();
let cache = self.introspect_cache.clone();
let share_id = share.id.clone();
let server = share.server.clone();
let export = share.export.clone();
let port = share.port;
let queued = work.len();
tokio::spawn(async move {
let mut upgraded = 0usize;
for job in work {
let probe = async {
let mut reader = NfsReadAt::open(&server, &export, port, &job.filename).await?;
let report =
introspect_reader(&mut reader, job.size, &job.filename, false).await;
reader.finish().await;
Ok::<_, NfsClientError>(report)
};
match tokio::time::timeout(INTROSPECT_TIMEOUT, probe).await {
Ok(Ok(report)) => {
cache.put(&share_id, &job.filename, job.size, report.clone());
if store.update_external_introspection(&job.iso_id, report) {
upgraded += 1;
}
}
Ok(Err(e)) => tracing::warn!(
target: "openpxe::nfs",
share = %share_id, iso = %job.filename,
"introspection failed: {e}"
),
Err(_) => tracing::warn!(
target: "openpxe::nfs",
share = %share_id, iso = %job.filename,
"introspection timed out after {}s", INTROSPECT_TIMEOUT.as_secs()
),
}
}
tracing::info!(
target: "openpxe::nfs",
share = %share_id, queued, upgraded,
"remote introspection pass complete"
);
});
}
/// v0.7.4: locate `in_iso_path` inside a share-hosted ISO. Returns the
/// byte range so the HTTP layer can serve kernel/initrd files out of
/// remote ISOs with a follow-up ranged [`Self::stream_iso`].
pub async fn locate_in_iso(
&self,
share_id: &str,
filename: &str,
in_iso_path: &str,
) -> Result<Option<FileLocation>> {
let share = self
.get(share_id)
.ok_or_else(|| Error::Invalid(format!("no such NFS share '{share_id}'")))?;
if filename.contains('/') || filename.contains('\\') || filename.contains("..") {
return Err(Error::Invalid(format!("invalid filename '{filename}'")));
}
let mut reader = NfsReadAt::open(&share.server, &share.export, share.port, filename)
.await
.map_err(|e| Error::Invalid(format!("nfs open '{filename}': {e}")))?;
let loc = iso_fs::lookup(&mut reader, in_iso_path).await;
reader.finish().await;
Ok(loc)
}
fn update_status(
&self,
id: &str,
@@ -490,6 +599,96 @@ struct NfsListEntry {
size: u64,
}
/// The connection type [`build_connection`] yields.
type NfsConn = nfs3_client::Nfs3Connection<nfs3_client::tokio::TokioIo<tokio::net::TcpStream>>;
/// v0.7.4: random-access reader over one NFS connection + file handle —
/// the [`IsoReadAt`] impl that lets the ISO9660 walker and introspection
/// probes run against share-hosted images.
struct NfsReadAt {
conn: NfsConn,
fh: nfs_fh3,
}
impl NfsReadAt {
/// Connect, mount, and LOOKUP `filename` at the export root.
async fn open(
server: &str,
export: &str,
port: u16,
filename: &str,
) -> std::result::Result<Self, NfsClientError> {
let mut conn = build_connection(server, export, port).await?;
let root = conn.root_nfs_fh3();
let lookup = conn
.lookup(&LOOKUP3args {
what: diropargs3 {
dir: root,
name: filename3(Opaque::borrowed(filename.as_bytes())),
},
})
.await
.map_err(NfsClientError::Rpc)?;
let fh = match lookup {
Nfs3Result::Ok(o) => o.object,
Nfs3Result::Err((status, _)) => {
return Err(NfsClientError::Nfsstat(status_label(status)));
}
};
Ok(Self { conn, fh })
}
/// Best-effort unmount. Consumes the reader — it's done.
async fn finish(self) {
let _ = self.conn.unmount().await;
}
}
impl IsoReadAt for NfsReadAt {
async fn read_at(&mut self, offset: u64, len: u32) -> std::io::Result<Vec<u8>> {
let mut out: Vec<u8> = Vec::with_capacity(len as usize);
let mut off = offset;
// READ3 may legally return fewer bytes than asked (server cap);
// loop until the exact-read contract is satisfied or the file
// genuinely ends short.
while (out.len() as u32) < len {
let want = (len - out.len() as u32).min(READ_CHUNK_BYTES);
let res = self
.conn
.read(&READ3args {
file: self.fh.clone(),
offset: off,
count: want,
})
.await
.map_err(|e| std::io::Error::other(e.to_string()))?;
let ok = match res {
Nfs3Result::Ok(o) => o,
Nfs3Result::Err((status, _)) => {
return Err(std::io::Error::other(status_label(status)));
}
};
let data = ok.data.as_ref();
if data.is_empty() {
return Err(std::io::Error::new(
std::io::ErrorKind::UnexpectedEof,
"NFS read past end of file",
));
}
out.extend_from_slice(data);
off += data.len() as u64;
if ok.eof && (out.len() as u32) < len {
return Err(std::io::Error::new(
std::io::ErrorKind::UnexpectedEof,
"NFS read past end of file",
));
}
}
out.truncate(len as usize);
Ok(out)
}
}
/// Connect, READDIR the export root, look up each `*.iso` to get its
/// size + file handle. Returns a flat list. Errors are returned with
/// a human-readable message; the caller decides how to surface them.