README gains a Library section covering prepare, the exec options (envp, bind targets, arch) and the provisioned-rootfs hot-cell flow. SPEC documents the workspace layout, the library API, the cache sidecars (.digest, .provisioned) and the updated execution flow. AGENTS.md scope paths follow the new crates/ layout.
400 lines
14 KiB
Markdown
400 lines
14 KiB
Markdown
# ecr - implementation specification
|
|
|
|
## Workspace Layout
|
|
|
|
Cargo workspace with two crates:
|
|
|
|
- `crates/ecr` — the library (package `ecr`). All functionality lives here:
|
|
image resolution, download, extraction, namespaces, mounts, chroot exec,
|
|
QEMU VM mode.
|
|
- `crates/ecr-cli` — the CLI front-end (package `ecr-cli`, binary `ecr`).
|
|
Parses flags, maps them onto the library API, propagates exit codes.
|
|
|
|
The library must not depend on CLI types (`clap` is a CLI-only dependency).
|
|
|
|
## Library API
|
|
|
|
### Rootfs lifecycle (`ecr::rootfs`)
|
|
|
|
- `RootfsCache::new(dir)` / `RootfsCache::default_dir()` — the image tarball
|
|
cache (default `~/.cache/ecr`).
|
|
- `cache.prepare(&PrepareRequest)` — parse the image reference, resolve
|
|
floating version aliases ("latest", "lts", "edge") to a concrete version
|
|
before computing the cache key, run the OCI `:latest` digest freshness
|
|
check, download through the cache when needed, and extract into a scratch
|
|
directory. Returns a `PreparedRootfs` (TempDir-backed; dropped on drop).
|
|
- `cache.prepare_provisioned(&req, provision)` — hot-cell amortization: on a
|
|
cache hit with a `.provisioned` sidecar the entry is extracted as-is (no
|
|
freshness check, no provisioning); otherwise the rootfs is prepared,
|
|
`provision(&PreparedRootfs)` runs once, and `persist` writes the result
|
|
back. `no_cache` forces re-download + re-provision.
|
|
- `PreparedRootfs::persist()` — packs the current rootfs into its cache
|
|
entry (compression chosen from the entry's filename extension: gz, xz,
|
|
zstd, plain tar; symlinks stored as links; permissions preserved) and
|
|
writes the `.provisioned` sidecar. The entry's contents are replaced:
|
|
plain `prepare` also returns the provisioned contents from then on.
|
|
|
|
### Execution (`ecr::exec`)
|
|
|
|
- `exec(rootfs, &ExecOptions)` — run a command inside a prepared rootfs in
|
|
fresh user/PID/mount/UTS namespaces; returns the exit code (128+signal on
|
|
kill). Empty option fields select defaults:
|
|
- `arch`: host architecture; a foreign arch requires binfmt_misc.
|
|
- `env`: caller-composed envp as `(key, value)` pairs; empty selects
|
|
`chroot::default_env`. The host environment is never inherited.
|
|
- `binds`: explicit `BindTarget { source, target, read_only }` — host
|
|
directory mounted at an absolute in-rootfs mount point; read-only via
|
|
overlay, read-write via bind. `..` targets are rejected.
|
|
- `dns`: nameservers written to /etc/resolv.conf; empty copies the host
|
|
resolver.
|
|
- `command`: argv; empty runs the rootfs default shell.
|
|
- `working_dir`: explicit in-rootfs cwd; empty picks the first read-write
|
|
bind target that exists, then `/root`, then `/`.
|
|
|
|
### Mount plumbing (`ecr::mount`)
|
|
|
|
`setup_mounts(rootfs, &[BindTarget])` mounts proc/dev/devpts/sys and applies
|
|
the bind targets; `in_rootfs(rootfs, target)` maps in-chroot paths and
|
|
rejects escaping targets. Overlay upper/work directories are temp dirs the
|
|
caller must keep alive (returned by `setup_mounts`).
|
|
|
|
## Synopsis
|
|
|
|
```
|
|
ecr [OPTIONS] <DISTRO[:VERSION]> -- [COMMAND]...
|
|
```
|
|
|
|
## CLI Interface
|
|
|
|
### Positional Arguments
|
|
|
|
- `<distro>` (required): Distribution name or OCI image reference
|
|
- `<version>` (optional): Distribution version/codename
|
|
|
|
### Options
|
|
|
|
| Flag | Default | Description |
|
|
|------|---------|-------------|
|
|
| `-a, --arch <arch>` | host arch | Target architecture |
|
|
| `--bind <path>` | cwd | Directory to overlay-mount (can be specified multiple times) |
|
|
| `--bind-rw <path>` | none | Read-write bind mount at `/mnt/<basename>` (can be specified multiple times, overrides `--bind` for same path) |
|
|
| `--no-cache` | false | Download fresh tarball, ignore cache |
|
|
| `--no-bind` | false | Skip mounting any directory |
|
|
| `--kernel[=PATH]` | none | Boot with QEMU system emulation; downloads the default Alpine `linux-virt` kernel when no `=PATH` is given |
|
|
| `-m, --memory <size>` | 2G | Memory size for QEMU VM (only used with `--kernel`) |
|
|
| `-v, --verbose` | false | Print diagnostic messages |
|
|
| `-h, --help` | - | Show help |
|
|
| `-V, --version` | - | Show version |
|
|
|
|
## File Layout
|
|
|
|
### Cache Directory
|
|
|
|
```
|
|
~/.cache/ecr/
|
|
├── ubuntu-noble-amd64.tar.gz
|
|
├── alpine-latest-x86_64.tar.gz
|
|
├── oci-docker_io-library_archlinux-latest-amd64.tar.gz
|
|
├── oci-docker_io-library_archlinux-latest-amd64.tar.gz.digest
|
|
└── ...
|
|
```
|
|
|
|
Sidecar files, never counted as image entries:
|
|
|
|
- `<entry>.digest` — manifest digest of the last OCI download, used by the
|
|
`:latest` freshness check.
|
|
- `<entry>.provisioned` — marker written by `PreparedRootfs::persist`;
|
|
`prepare_provisioned` treats an entry with this marker as provisioned.
|
|
|
|
Tarballs are downloaded once and never redownloaded (unless the digest
|
|
moves, `--no-cache` is passed, or a provisioned entry is deleted). Users
|
|
can delete files manually.
|
|
|
|
### Config File
|
|
|
|
`~/.config/ecr.yaml`:
|
|
|
|
```yaml
|
|
dns:
|
|
- 1.1.1.1
|
|
```
|
|
|
|
## Distro Sources
|
|
|
|
### Direct Tarball Downloads
|
|
|
|
| Distro | Version Format | Source |
|
|
|--------|----------------|--------|
|
|
| Ubuntu | noble, jammy, mantic or 26.04, 25.10, 22.04, latest, lts | cdimage.ubuntu.com |
|
|
| Alpine | 3.20, 3.19, latest, edge | dl-cdn.alpinelinux.org |
|
|
|
|
### Docker Hub (OCI Registry)
|
|
|
|
All other distributions use Docker Hub images via OCI registry API:
|
|
|
|
| Distro | Image Reference |
|
|
|--------|-----------------|
|
|
| Debian | `library/debian` |
|
|
| Arch | `library/archlinux` |
|
|
| Fedora | `library/fedora` |
|
|
| Gentoo | `gentoo/stage3` |
|
|
| Custom | `<image>[:tag]` or `<registry>/<image>[:tag]` |
|
|
|
|
### Custom Image References
|
|
|
|
Users can specify any OCI-compatible image:
|
|
|
|
```
|
|
ecr debian:bookworm -- ./build.sh
|
|
ecr gentoo/stage3 -- emerge --sync
|
|
ecr gcr.io/my-project/my-image:v1.0 -- /app/test
|
|
```
|
|
|
|
### Architecture Mapping
|
|
|
|
| ecr | Ubuntu | Alpine | Docker Hub |
|
|
|-----|--------|--------|------------|
|
|
| amd64 | amd64 | x86_64 | amd64 |
|
|
| arm64 | arm64 | aarch64 | arm64 |
|
|
| armhf | armhf | armv7 | arm/v7 |
|
|
| riscv64 | riscv64 | riscv64 | riscv64 |
|
|
| ppc64el | ppc64el | ppc64le | ppc64le |
|
|
| s390x | s390x | s390x | s390x |
|
|
|
|
### OCI Image Download
|
|
|
|
For Docker Hub images:
|
|
|
|
1. Get anonymous bearer token from `https://auth.docker.io/token`
|
|
2. Query manifest list: `GET https://registry.hub.docker.com/v2/<repo>/manifests/<tag>`
|
|
3. Select manifest matching target architecture
|
|
4. Download layer blobs with authentication
|
|
5. Extract layers to rootfs
|
|
|
|
If architecture is not available in manifest list, error with available architectures:
|
|
|
|
```
|
|
Error: No manifest found for architecture 'riscv64'. Available: amd64, arm64, ppc64le, s390x
|
|
```
|
|
|
|
## Execution Flow
|
|
|
|
The CLI delegates to the library; the namespace-mode flow is:
|
|
|
|
1. Parse CLI arguments, map flags onto library requests
|
|
2. `cache.prepare`: resolve distro/version/arch to image source
|
|
3. Check cache for existing tarball
|
|
4. If not cached, download tarball (direct or OCI)
|
|
5. Extract tarball to a temporary directory
|
|
6. `ecr::exec`: create namespaces: user, pid, mount, uts
|
|
7. Set up mounts: /proc, /sys (ro), /dev, /dev/pts
|
|
8. Apply bind targets: overlays (ro) and bind mounts (rw)
|
|
9. Write /etc/resolv.conf with DNS servers
|
|
10. Set the working directory
|
|
11. Exec shell or command in chroot with the composed envp
|
|
12. On exit, clean up the temporary directory; propagate the exit code
|
|
|
|
## Namespace Setup
|
|
|
|
### Namespaces (Always Created)
|
|
|
|
- **user**: Map current user to root (UID 0) inside
|
|
- **pid**: Isolated process tree
|
|
- **mount**: Private mounts for chroot setup
|
|
- **uts**: Hostname set to `ecr-<distro>-<random>`
|
|
|
|
### Network
|
|
|
|
Host network namespace (no isolation).
|
|
|
|
### User Namespace Mapping
|
|
|
|
```
|
|
uid_map: 0 <current_uid> 1
|
|
gid_map: 0 <current_gid> 1
|
|
```
|
|
|
|
This makes the user appear as root inside the chroot while remaining unprivileged on the host.
|
|
|
|
### Mounts Inside Chroot
|
|
|
|
| Path | Type | Options |
|
|
|------|------|---------|
|
|
| /proc | proc | defaults |
|
|
| /sys | sysfs | ro,nosuid,nodev,noexec |
|
|
| /dev | devtmpfs | nosuid |
|
|
| /dev/pts | devpts | nosuid,noexec |
|
|
| /root/<basename> | overlay | lowerdir=<bind_path>, upperdir=<temp>, workdir=<temp> |
|
|
| /mnt/<basename> | bind | rw (for --bind-rw) |
|
|
| /etc/resolv.conf | file | written with DNS |
|
|
|
|
## QEMU Integration
|
|
|
|
### Foreign Architecture Detection
|
|
|
|
If `--arch` differs from host architecture, QEMU is required.
|
|
|
|
### binfmt_misc Check
|
|
|
|
Before entering chroot, verify binfmt_misc is registered for target architecture by checking `/proc/sys/fs/binfmt_misc/qemu-<arch>`.
|
|
|
|
If not registered, error with message:
|
|
|
|
```
|
|
Error: binfmt_misc not registered for riscv64
|
|
|
|
Install QEMU user emulation:
|
|
Ubuntu/Debian: sudo apt install qemu-user-static
|
|
Arch: sudo pacman -S qemu-user-static-binfmt
|
|
Alpine: sudo apk add qemu-user-static
|
|
```
|
|
|
|
### QEMU Binary
|
|
|
|
No action required. Modern qemu-user-static packages register binfmt_misc with the `F` (fix binary) flag, loading the interpreter into kernel memory. The kernel handles foreign binary execution transparently.
|
|
|
|
## QEMU System Emulation Mode
|
|
|
|
When `--kernel` is specified, ecr switches from namespace/chroot mode to QEMU system emulation. The extracted rootfs is converted to an uncompressed CPIO initramfs and booted with the provided kernel.
|
|
|
|
### Usage
|
|
|
|
The kernel path uses `=` syntax (`--kernel=PATH`); `--kernel` without a value downloads the default kernel. Without `=`, a following path would be parsed as the DISTRO argument.
|
|
|
|
```sh
|
|
ecr --kernel ubuntu:noble
|
|
ecr --kernel=/boot/vmlinuz ubuntu:noble
|
|
ecr --kernel=/boot/vmlinuz --memory 4G alpine
|
|
ecr --kernel=/boot/vmlinuz debian -- /bin/sh -c "echo hello"
|
|
```
|
|
|
|
### Execution Flow
|
|
|
|
1. Download/cache rootfs tarball (same as namespace mode)
|
|
2. Extract tarball to temporary directory
|
|
3. Create uncompressed CPIO initramfs from rootfs (streamed to disk), including essential device nodes (/dev/ttyS0, /dev/null, /dev/tty) and an `/init` script that mounts proc/sys/dev, sets the hostname, execs the requested command argv verbatim (each argv element base64-encoded in the cmdline as `ECR_ARGV`), and powers off on exit
|
|
4. Launch QEMU with:
|
|
- `-kernel <path>` - provided (or downloaded) kernel
|
|
- `-initrd initramfs.cpio` - rootfs as initramfs
|
|
- `-append "console=ttyS0 [quiet] ECR_SHELL=... [ECR_ARGV=...] ECR_HOSTNAME=..."` - kernel command line (`quiet` unless `-v`)
|
|
- `-m <memory>` - memory size (default 2G)
|
|
- `-display none -serial mon:stdio` - console on stdio
|
|
- `-netdev user,id=net0 -device virtio-net-pci,netdev=net0` - network NIC
|
|
- `-enable-kvm -cpu host` - when the host supports KVM and the target matches the host architecture
|
|
5. Wait for QEMU to exit (init powers the VM off when the command/shell exits; `-no-reboot` makes QEMU terminate)
|
|
6. Cleanup temporary files
|
|
|
|
### Initramfs Creation
|
|
|
|
The rootfs directory is converted to an uncompressed CPIO archive (newc format) using the `cpio` crate, streamed entry by entry so large rootfs images never need to fit in memory. Hard links are preserved: the first occurrence of a (device, inode) pair carries the data with a synthetic inode, subsequent occurrences are zero-size entries sharing that inode, which the kernel's initramfs loader turns into real hard links.
|
|
|
|
### Architecture Support
|
|
|
|
| ecr Arch | QEMU System Binary |
|
|
|----------|-------------------|
|
|
| amd64/x86_64 | qemu-system-x86_64 |
|
|
| arm64/aarch64 | qemu-system-aarch64 |
|
|
| armhf/armv7 | qemu-system-arm |
|
|
| riscv64 | qemu-system-riscv64 |
|
|
| ppc64el | qemu-system-ppc64 |
|
|
| s390x | qemu-system-s390x |
|
|
|
|
### Requirements
|
|
|
|
- QEMU system emulator installed (`qemu-system-<arch>`)
|
|
- Kernel with required drivers (serial console, virtio-net for network)
|
|
|
|
### Differences from Namespace Mode
|
|
|
|
| Feature | Namespace Mode | QEMU Mode |
|
|
|---------|---------------|-----------|
|
|
| Isolation | User namespace | Full VM |
|
|
| Performance | Near-native | Emulated (KVM-accelerated when available) |
|
|
| Root access | No | Yes (inside the VM) |
|
|
| Foreign arch | binfmt_misc required | Built-in emulation |
|
|
| Bind mounts | Overlay/bind | Not supported (flags are ignored with a warning) |
|
|
| Network | Host network | User-mode NIC (not configured inside the guest) |
|
|
|
|
## File Handling
|
|
|
|
### Overlay Mount (Default)
|
|
|
|
By default, the current working directory is mounted as an overlay filesystem at `/root/<basename>` inside the chroot, where `<basename>` is the name of the current directory.
|
|
|
|
Overlay configuration:
|
|
- `lowerdir`: the source directory (read-only)
|
|
- `upperdir`: temp directory for modifications
|
|
- `workdir`: temp directory required by overlayfs
|
|
|
|
Changes made inside the chroot are written to upperdir and discarded on exit. The host directory is never modified.
|
|
|
|
Multiple `--bind` paths can be specified, each creates an overlay at `/root/<basename>`.
|
|
|
|
Example:
|
|
```
|
|
$ cd ~/projects/myapp
|
|
$ ecr ubuntu:noble -- make build
|
|
# ~/projects/myapp mounted at /root/myapp
|
|
# Build artifacts written to overlay, discarded on exit
|
|
```
|
|
|
|
### Read-Write Bind Mount
|
|
|
|
`--bind-rw <path>` creates a true read-write bind mount at `/mnt/<basename>`. This modifies the host filesystem directly. Use with caution.
|
|
|
|
Multiple `--bind-rw` paths can be specified. If a path is specified in both `--bind` and `--bind-rw`, the read-write mount takes precedence.
|
|
|
|
If no path is specified, defaults to current working directory.
|
|
|
|
### No Mount
|
|
|
|
`--no-bind` skips mounting any directory.
|
|
|
|
## DNS
|
|
|
|
Default DNS server is 1.1.1.1. Configured via `/etc/resolv.conf` in chroot:
|
|
|
|
```
|
|
nameserver 1.1.1.1
|
|
```
|
|
|
|
Override with config file (`~/.config/ecr.yaml`):
|
|
|
|
```yaml
|
|
dns:
|
|
- 8.8.8.8
|
|
- 8.8.4.4
|
|
```
|
|
|
|
## Environment Variables
|
|
|
|
Default environment inside chroot (`chroot::default_env`, used by the CLI;
|
|
library callers compose their own envp):
|
|
|
|
- HOME=/root
|
|
- USER=root
|
|
- SHELL=/bin/bash (or /bin/sh if bash unavailable)
|
|
- TERM=<from host>
|
|
- PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
|
|
|
|
Host environment is not inherited.
|
|
|
|
## Signal Handling
|
|
|
|
Forward SIGINT, SIGTERM, SIGHUP, SIGQUIT to child process. Wait for child to exit before cleanup.
|
|
|
|
## Security Requirements
|
|
|
|
### User Namespace Required
|
|
|
|
`ecr` requires unprivileged user namespaces. If unavailable (sysctl `kernel.unprivileged_userns_clone=0` or AppArmor restrictions), error with:
|
|
|
|
```
|
|
Error: User namespaces not available
|
|
|
|
Enable with:
|
|
sysctl -w kernel.unprivileged_userns_clone=1
|
|
|
|
Or check AppArmor profile restrictions.
|
|
```
|