Architecture Decisions

Key architectural decisions for xSwarm dashboard and security infrastructure

This document records significant architectural decisions for the xSwarm platform.


ADR-001: Dashboard Restructure - Project-Centric Architecture

Status: Planned Date: January 2026

Context

The current dashboard has a flat, task-centric view where tasks appear at the global level. This doesn’t match the mental model where tasks belong to projects, and projects are assigned to worker machines.

Decision

Transform the dashboard from flat task-centric view to hierarchical project-centric architecture with worker capacity management and console access.

Key Changes

  • Tasks belong to projects (not shown at global level)
  • Projects get rich dashboards with activity, chat, console
  • Workers show capacity metrics and project assignments
  • Project “move” = clone repo + decrypt secrets + npm install + assign worker

Database Schema Updates

Workers Table Additions

disk_total: integer('disk_total'),      // MB
disk_free: integer('disk_free'),        // MB
ram_total: integer('ram_total'),        // MB
cpu_cores: integer('cpu_cores'),
gpu_name: text('gpu_name'),
gpu_vram: integer('gpu_vram'),          // MB
tasks_completed: integer('tasks_completed').default(0),
hours_worked: real('hours_worked').default(0),

Projects Table Additions

assigned_worker_id: text('assigned_worker_id').references(() => workers.id, { onDelete: 'set null' }),
color: text('color'),                   // Unique hex color
last_activity_at: text('last_activity_at'),

New Tables

  • chat_messages - Global and project-specific chat
  • activity_log - Project activity timeline

API Endpoints

Workers Routes

  • POST /workers/:id/heartbeat - Add capacity metrics
  • GET /workers/:id/projects - Get assigned projects

Projects Routes

  • POST /projects/:id/assign - Assign project to worker
  • GET /projects/:id/activity - Get activity timeline

New Chat Routes

  • GET /chat?project_id=xxx - Get messages
  • POST /chat - Send message
  • GET /notifications - Cross-project notifications

UI Changes

Main Dashboard

  • Remove global tasks section
  • Add project cards grid with unique colors, activity indicators
  • Add quick stats: total projects, active workers, pending issues
  • Add global chat panel (right sidebar) with project switcher
  • Add notifications feed from all projects

Per-Project Dashboard

Tabs:

  1. Overview - Header, activity timeline, current tasks, kanban link
  2. Tasks - Existing kanban view
  3. Chat - Project-specific chat panel
  4. Console - Terminal access to worker (xterm.js)
  5. Settings - Worker assignment, secrets management

Workers Page

Display per worker:

  • Capacity: CPU cores, RAM (used/total), Disk (used/total), GPU
  • Assigned Projects: List with colors
  • Activity: Tasks completed, hours worked
  • Actions: “Move Project” button

Project-Worker Assignment Flow

  1. Update projects.assigned_worker_id in DB
  2. Worker detects assignment (polling)
  3. Worker clones GitHub repo
  4. Worker fetches + decrypts secrets
  5. Worker runs install scripts
  6. Worker reports success/failure

Console Access Architecture (xterm.js)

  1. Dashboard opens xterm.js terminal
  2. WebSocket to API → routes to worker
  3. Worker spawns PTY, bridges to WebSocket

Implementation Order

  1. Sprint 1: Schema + API (Database, endpoints)
  2. Sprint 2: Main dashboard + Workers page
  3. Sprint 3: Per-project dashboard with tabs
  4. Sprint 4: Assignment flow + Console access

Consequences

  • More intuitive navigation aligned with mental model
  • Better visibility into worker capacity
  • Enables future automatic load balancing
  • Requires migration of existing data

ADR-002: Zero-Knowledge User Data Encryption

Status: Planned Date: January 2026

Context

Users store sensitive data in xSwarm: BYOK API keys (OpenAI, Anthropic), project secrets (.env files), planning discussions. The platform should not have access to this data under any circumstances.

Decision

Implement true zero-knowledge encryption where all encryption/decryption happens on the client (CLI or browser). Server only stores encrypted blobs.

Requirements

  • Only Requirement: GitHub account (already required for xSwarm)
  • No password management: Use device approval model (like Signal/WhatsApp)
  • Recovery: Single recovery code for disaster scenarios

Architecture Overview

┌─────────────────────────────────────────────────────────────┐
│                  ADMIN DATABASE (xSwarm-readable)           │
│  - User auth (GitHub ID, email)                             │
│  - Billing (Stripe ID, subscription)                        │
│  - Project metadata (IDs, task counts for billing)          │
│  - Device registry (public keys for key transfer)           │
│  - Pending approvals (encrypted key blobs in transit)       │
└─────────────────────────────────────────────────────────────┘
                              │
                              │ User DB URL (server can connect
                              │ but only sees encrypted blobs)
                              ▼
┌─────────────────────────────────────────────────────────────┐
│              USER PERSONAL DATABASE (encrypted)             │
│  - Projects (name, description, repo_url encrypted)         │
│  - Tasks (title, description, criteria encrypted)           │
│  - BYOK API keys (OpenAI/Anthropic keys encrypted)          │
│  - Project secrets (.env files encrypted)                   │
│  - Planning meetings (specs, decisions encrypted)           │
└─────────────────────────────────────────────────────────────┘

Key Management: Device Approval Model

First Device Ever

  1. GitHub OAuth login
  2. System generates master encryption key
  3. System generates recovery code: “ALPHA-BRAVO-CHARLIE-DELTA”
  4. Display: “Save this recovery code somewhere safe”
  5. User confirms (checkbox)
  6. Key stored locally (~/.xswarm/key or IndexedDB)

New Device (Have Existing Device)

  1. GitHub OAuth login on new device
  2. New device generates keypair, registers public key
  3. “Waiting for approval from existing device…”
  4. Existing device shows: “New device requesting access [Approve] [Deny]”
  5. User clicks Approve
  6. Existing device encrypts master key with new device’s public key
  7. Encrypted blob sent via server (server can’t decrypt)
  8. New device decrypts, stores key locally

Disaster Recovery (Lost All Devices)

  1. GitHub OAuth login
  2. “No approved devices found. Enter recovery code:”
  3. User enters recovery code
  4. Master key derived from recovery code (PBKDF2)
  5. Key stored locally, device registered as approved

Browser ↔ CLI (Same Machine)

  1. CLI runs localhost HTTP server on port 19284
  2. Browser detects localhost server
  3. Browser requests key transfer
  4. CLI prompts user: “Browser requesting access [Allow]”
  5. User allows → key sent over localhost (never hits internet)

Encryption Details

Algorithm: AES-256-GCM (WebCrypto API - native in browser + Node.js) Encrypted data format:

{
  "iv": "<base64-96-bit-random>",
  "ciphertext": "<base64-encrypted-data>",
  "version": 1
}

Database Schema Changes

Admin DB Additions

// Users table additions
user_db_url: text('user_db_url'),              // database URL for user's encrypted DB
user_db_token: text('user_db_token'),          // Auth token for user DB
recovery_code_hash: text('recovery_code_hash'), // bcrypt hash (to verify, not derive)
encryption_version: integer('encryption_version').default(1),
// New table: devices (for device approval flow)
devices: {
  id, user_id,
  name: text,                    // "Chad's MacBook Pro"
  device_type: text,             // "cli" | "web"
  public_key: text,              // For end-to-end key transfer
  created_at, last_seen_at,
  is_approved: boolean
}
// New table: pending_approvals (temporary, during key transfer)
pending_approvals: {
  id, user_id,
  requesting_device_id,
  encrypted_key_blob: text,      // Encrypted with requester's public key
  created_at, expires_at         // Auto-expire after 10 minutes
}
// New table: projects_metadata (for billing - no sensitive data)
projects_metadata: {
  id, user_id, slug, status,
  task_count, last_activity_at,
  kanban_token, kanban_enabled
}

User DB Schema

// All *_encrypted fields store JSON: {iv, ciphertext, version}
projects: { id, slug, name_encrypted, description_encrypted, ... }
tasks: { id, project_id, status, title_encrypted, description_encrypted, ... }
byok_api_keys: { id, provider, key_encrypted, name_encrypted, ... }
project_secrets: { id, project_id, file_path, content_encrypted }
planning_meetings: { id, project_id, status, title_encrypted, ... }

Implementation Files

New Files

File Purpose
packages/shared/crypto.js Isomorphic encrypt/decrypt, keypair generation
packages/shared/user-db-schema.js User DB Drizzle schema
packages/shared/wordlist.js Recovery code word list (BIP39 subset)
packages/api/src/db/schema.js devices, pending_approvals tables
packages/api/src/services/user-db.js Per-user DB provisioning via D1 API
packages/api/src/routes/devices.js Device registration, approval endpoints
packages/web/src/stores/crypto.js Svelte store for encryption key
packages/web/src/components/DeviceApproval.svelte Approval UI
packages/app/src/daemon/crypto.js CLI encryption + localhost server

Modified Files

File Changes
packages/api/src/db/schema.js Add devices, pending_approvals, projects_metadata
packages/api/src/routes/auth.js Device registration on login
packages/app/src/daemon/auth.js Store/load encryption key, run localhost server
packages/app/src/index.js Handle first-time setup flow
packages/web/src/stores/auth.js Integrate device approval flow

Security Considerations

  1. Master key never sent to server - only encrypted blobs transit
  2. Device keypairs for transfer - asymmetric encryption prevents server access
  3. Recovery code → key derivation - PBKDF2 with 310k iterations
  4. Unique IV per encryption - prevents pattern analysis
  5. Authenticated encryption (GCM) - detects tampering
  6. Pending approvals expire - 10 minute window limits exposure
  7. Localhost-only browser transfer - never hits internet
  8. No recovery without code - true zero-knowledge trade-off

Implementation Phases

  1. Phase 1: Crypto Module - AES-256-GCM, keypair generation, recovery codes
  2. Phase 2: Database Schema - devices, pending_approvals, user DB schema
  3. Phase 3: Device API - registration, approval, key transfer endpoints
  4. Phase 4: CLI Integration - setup flow, localhost server, encryption wrapper
  5. Phase 5: Web Integration - device approval UI, IndexedDB storage
  6. Phase 6: Testing - unit tests, integration tests, security audit

Consequences

  • True zero-knowledge: xSwarm cannot access user data
  • Frictionless day-to-day: no passwords, just device approval
  • Disaster recovery: single code to save
  • Complexity: multi-device key synchronization
  • Migration: existing users need onboarding flow

The pipeline: context for review

ADR-004 through ADR-013 describe the autonomous delivery pipeline. They are written to be read without the codebase, because their purpose is external review. This section is the system model they assume.

What the system does

A ticket carrying executable acceptance criteria is picked up by a headless coding agent running on the user’s own machine, in a git worktree of their repository. The agent writes a failing test, implements it, and iterates until green. The change is then cleaned up, merged to trunk behind a tested merge candidate, and deployed — with no human reviewing the code at any point. The human files the idea and approves the result.

The constraints that shaped it

  • Execution is local. Source never leaves the user’s machine. The hosted API carries tickets down and status up.
  • The repository is private, on a free plan. No paid GitHub features. Branch protection and rulesets appear available but are unverified; nothing depends on them.
  • No CI service. Deployment is a push to staging or production, which the host already subscribes to. GitHub Actions would cost 2–3 minutes per run against a 2,000-minute monthly allowance, duplicating a suite that must run locally anyway.
  • Agents are unreliable narrators. Every design choice below assumes an agent may report success it did not achieve — not from malice, but because “I think it works” is the cheapest output it can produce.

The failure these decisions were written against

On 2026-09-21 the board reported 91 tickets deployed to production. Measured against git and the live service: 77 had never run an agent, 80 carried no commit sha, 27 shared a single landedSha belonging to a different ticket’s merge commit, and 96 rows were marked completed against only 24 recorded completion events. Production had not moved in days.

The cause was two functions that ran before every dispatch and advanced tickets with no agent involved. Nothing was lying; the system was simply asking a question (“is this sha an ancestor of what production serves?”) whose answer is true for reasons unrelated to the ticket.

Every ADR below is a consequence of that, and the recurring principle is: a claim is not evidence, and absence of a check is not a pass.

The vocabulary

Term Meaning
ticket a unit of work with executable acceptance criteria (Given/When/Then) and a named test path
stage one of six: Planning, Coding, Testing, Refactoring, Staging, Deployed
evidence named boolean keys recorded by the step that performed the work — testObservedFailing, commitExists, deployVerified, …
gate something that refuses. Distinct from a report, which informs and cannot block
merge queue builds a candidate (trunk + the change), tests it, and fast-forwards trunk only if green
unearned advance a ticket sitting past a stage whose proof is incomplete — the shape the 91 took

What would be most useful to review

  1. ADR-011 (squash vs fast-forward onto main) — explicitly unresolved, and the decision I am least confident in.
  2. ADR-009’s residual risk. Gates run inside the pipeline, so if the pipeline’s own gates are wrong nothing external catches it. Required status checks would move enforcement outside, at the cost of a second authority that could disagree. Is that trade worth taking?
  3. ADR-010’s enforceability. The agent is asked to commit each iteration and the result is measured, but not enforced. Is measured-not-enforced sufficient, or does a non-complying agent need to fail the ticket?
  4. ADR-006’s asymmetry. Planning is scored but cannot block. Is a stage that can never refuse actually a stage?
  5. Whether the evidence model has a gap — a proof key that could be recorded truthfully while the underlying property is false.

ADR-004: Stages advance on recorded evidence, never on a report

Status: Implemented Date: September 2026

Context

An agent that reports success is the cheapest possible signal and the least reliable. On 2026-09-21 the board showed 91 tickets “deployed”: 77 had never run an agent, 80 had no commit sha, and 27 shared one landedSha belonging to a different ticket’s merge commit. Two functions running before every dispatch had written the flags — sweepDeployed granted deployVerified from an ancestry check, settleTerminal set completed by reading stored evidence back.

Decision

A stage advances only when its required evidence keys are present and literally true. Evidence is written by the step that performed the work, never by a sweep, and never by a later pass inferring it.

Alternatives rejected

  • Trust the agent’s report. This is what produced the 91.
  • Infer from git ancestry. Monotone in the deploy, not in the ticket: once production moves past a sha, every older sha qualifies for ever.

Consequences

Absence blocks. A ticket whose observer could not run stays put, which is correct but means an unwired observer silently stalls work — mitigated by recording a reason on every refusal.


ADR-005: Derive, never store

Status: Implemented Date: September 2026

Context

Every stored copy of a computed fact is a clock someone must wind. The stage column, the board column and the gate’s answer were three copies of one question and drifted until 44 shipped tickets were filed under “Coding”.

Decision

Anything computable at read time is computed at read time: a ticket’s stage, its score, the release that shipped it (from the git tag containing its sha), and its pull request (from the branch). None are persisted.

Alternatives rejected

  • Materialise scores into a table. Fast, and wrong in exactly the way this system keeps failing: a score that outlives the facts behind it is another fabricated verdict.

Consequences

Reporting costs a computation and sometimes a network call. Back-scoring history is free, which is how 201 releases could be scored retroactively with no migration.


ADR-006: Six stages, with Planning scored rather than gated

Status: Implemented Date: September 2026

Context

Four stages (Coding, Refactoring, Merging, Deployed) put the code⇄test loop inside “Coding”, so a ticket that wrote code and a ticket that proved it were indistinguishable.

Decision

Planning, Coding, Testing, Refactoring, Staging, Deployed. Coding carries the commit; Testing carries the red/green pair and the E2E verdict. Planning carries no proof key — its quality is scored from the ticket’s shape (executable spec, named test, declared touches, design brief).

Alternatives rejected

  • Give Planning a required key (designed). Tried and reverted the same hour: stageFor walks to the first incomplete stage, so a key there outranks every later fact and re-filed all 191 tickets under Planning — including 91 serving in production.

Consequences

The split immediately relocated 51 “unearned” advances from Coding to Testing, which is where they actually were. Planning cannot block a ticket, by design.


ADR-007: One ticket at a time, enforced by a lock file

Status: Implemented Date: September 2026

Context

Concurrency was a cap = 1 parameter. A second xswarm run started by hand put two agents on one checkout with no warning.

Decision

<repo>/.xswarm/local/run.lock holds {pid, ticket, startedAt}. A second run refuses, naming the holder. A lock whose pid is not alive is reclaimed and its ticket returned to pending.

Alternatives rejected

  • Rely on the cap. A parameter is a promise; a file whose holder must be alive is a fact.
  • Reclaim stale tickets inside the tick loop. It cannot tell a dead run from a working one and would reset a ticket an agent is mid-way through.

Consequences

Fails open and says so if the lock cannot be written — a guard that cannot answer must not become a new way to stop the queue.


ADR-008: Execution is local; the API is a relay

Status: Implemented Date: September 2026

Context

Running agents in a hosted environment would mean shipping customers’ source to a third party.

Decision

Agents run on the user’s machine in a git worktree per ticket. The API carries tickets down and status up; it never holds code. The board is a view.

Consequences

Onboarding is npx xswarm login, not a repository grant. The pipeline depends on the user’s machine being up, and “the daemon is not running” becomes a first-class observable.


ADR-009: GitHub is a record, not a gate; deployment is a branch push

Status: Implemented Date: September 2026

Context

The obvious move is GitHub Actions for CI and required status checks for enforcement.

Decision

No Actions. Gates run locally and are authoritative. Deployment is a push to staging or production, which the host (Cloudflare) already subscribes to. GitHub carries the record — the pull request, the development history, the release tag.

Alternatives rejected

  • Run the gate in Actions. The quality gate takes 2–3 minutes; thirty runs a day exceeds the free private-repo allowance inside a month, and it duplicates a suite that must run locally anyway before a push is permitted.
  • Required status checks as the enforcement point. Plausible, and worth revisiting: it would put enforcement outside the pipeline’s own code, which matters because the failure in ADR-004 was the pipeline grading its own homework. Not adopted because the local gates already refuse and adding a second authority risks them disagreeing.

Consequences

Nothing enforces the model outside the pipeline. If the pipeline’s own gates are wrong, nothing catches it — the residual risk in ADR-004.


ADR-010: Commit every iteration, test first, and verify the trail

Status: Implemented Date: September 2026

Context

The prompt said “Do NOT commit”. The whole code-to-green loop happened inside one agent session and reached the branch as a single diff, discarding the trail of what was tried and what failed.

Decision

The agent’s first commit is the acceptance criterion made executable, run, and watched to FAIL — committed alone as test: with Tests: red. Then one commit per attempt, each labelled red or green, until green. Then push. The pipeline reads the commits back and records iterations, iterationsRed, iterationsUnlabelled and testCommittedFirst as evidence.

Alternatives rejected

  • Split the finished diff into three commits. Built first, then discarded: it is presentational. A red commit produced by slicing a completed change asserts something it never earned.
  • Drive the loop from the pipeline, calling the agent once per iteration. Verifiable, but pays agent startup and context cost per iteration.

Consequences

Compliance is measured, not assumed, but an agent that ignores the instruction still produces work — the contract is recorded as unmet rather than enforced.


ADR-011: Squash onto main; the trail lives on the branch

Status: Proposed — not implemented Date: September 2026

Context

ADR-010 puts failing commits on the branch deliberately. The merge queue currently fast-forwards, so those commits would land on main. Every sha on main is potentially a release.

Decision (proposed)

Squash each ticket onto main as one green, deployable commit carrying the evidence trailers. The iteration history stays on the branch and the pull request, which GitHub retains after merge.

Alternatives

  • Fast-forward everything (current). Preserves the full trail in main, at the cost of shas on the deployable branch where tests fail — which breaks bisect and could ship broken.
  • Rebase-and-merge. Keeps commits and linearity but rewrites shas, breaking landedSha.

Open question for review

Is one green commit per ticket on main the right trade against losing the trail from main itself? This is the decision I am least certain of.


ADR-012: Nothing outside the repository

Status: Partially implemented Date: September 2026

Context

Pipeline state was scattered across ~/.xswarm (24 entries shared by every checkout), ~/.cache/xswarm-*, /tmp profiles, and a sibling ~/Projects/.worktrees holding 2,616 directories and 9.8GB. One project’s audit run aged out another’s record; one run’s reaper killed another run’s browser.

Decision

Everything the pipeline writes lives under <repo>/.xswarm/local/. Ownership becomes a path prefix rather than a guess about a process. A ratchet test fails the build if a new file reaches for $HOME.

Status

The reaper is scoped and the ratchet is in place at 48 known offenders. The profile creators still write to os.tmpdir(), so the reaper currently collects nothing — it fails closed, which leaks but cannot harm another project.


ADR-013: The development record is published where a developer would look

Status: Implemented Date: September 2026

Context

Every fact about how a ticket was built — each attempt, why it failed, what the refactor touched, the cost — was recorded in a local SQLite file and read back by nothing.

Decision

xswarm publish writes the spec, an evidence checklist, one comment per attempt including the failures, the refactor summary and the scorecard to the ticket’s pull request. Idempotent by marker, so re-running posts only what is missing.

Alternatives rejected

  • Backfill history as pull requests. Impossible: a branch already merged into main has no diff, and GitHub answers “No commits between”. Landed tickets get their record on the commit instead — free, retroactive, attached to what shipped.
  • GitHub Issues as the ticket source. Rejected: two models of one reality is the defect this system keeps paying for, and the board is the product. Issue Forms as a one-way intake adapter remains attractive and unbuilt.

Verification Checklists

Dashboard Restructure

  • Build passes: npm run build
  • Project cards render with colors
  • Workers show capacity metrics
  • Project assignment works
  • Console tab opens terminal

Zero-Knowledge Encryption

  • First device setup shows recovery code
  • Recovery code can restore access (all devices lost)
  • New device can be approved from existing device
  • Key transfer is end-to-end encrypted (verify server logs)
  • Data in User DB is encrypted (verify via direct DB query)
  • Wrong recovery code fails gracefully
  • CLI and web can share key via localhost
  • Revoking a device removes its access
  • CLI and web use same encryption (interoperable)

ADR-003: Terminal Chrome is Punctuation, Not Paper

Status: Accepted Date: September 2026

Context

Chad’s third visual-quality complaint in one day: “the projects list still looks not nice. We are overusing the green cli box everywhere.” The .terminal-window treatment (glowing green border, title bar, red/yellow/green traffic-light dots) had become the default wrapper for almost anything — a stat tile, a single list row, a whole panel — instead of a deliberate accent. An audit of every .terminal-window usage across packages/web/src found the pattern applied at three different scales, only one of which is the actual terminal metaphor doing real work:

  • The metaphor, legitimately: the TerminalDemo.astro component and a genuine live-activity feed (ProjectActivityFeed.svelte) — content that either is a terminal or is a log. (This entry originally also cited TerminalHero.astro as the homepage hero. It had already stopped being the homepage hero by the time this was written, and nothing imported it; it was deleted on 2026-09-10. A 2687-line component that reads as load-bearing and is reachable from nothing is worse than no component at all.)
  • A single panel wrapper: settings pages, login, most per-project side panels (ProjectBlockers.svelte, ProjectKanban.svelte, ProjectSessionsPanel.svelte). Defensible on its own, but several of these pages stack four or five of them, which is the “more than one boxed element” failure even when no single instance looks wrong.
  • Per-row/per-item, repeated: ProjectStats.svelte (five KPI tiles, each its own glowing box), ProjectGrid.svelte and dashboard/projects/index.astro (one fake terminal window per project card — 42 of them on the real page), FleetActivity.svelte and RequestsBoard.svelte (one per activity/request row), dashboard/workers/index.astro and dashboard/documents/index.astro (one per worker/document row). This is the concrete defect: a data-dense list where every row repeats the full chrome — border, title bar, three decorative dots that carry no status — competing with the data instead of presenting it. The traffic-light dots in particular are decoration pretending to be status in a product where colour is supposed to mean something real.

Decision

The terminal box is punctuation, not paper. Reserve .terminal-window for places where the terminal metaphor is doing real work — a hero, a footer moment, a code block, an actual live log. It is not the default container for a card, a row, or a panel.

Data-dense surfaces get hierarchy from typography and space, not borders. A list or table of many similar items (projects, workers, documents, requests, activity) uses hairline row separators, generous vertical rhythm, aligned columns with tabular numerals (font-variant-numeric: tabular-nums), and type weight/size to carry hierarchy — not a box per row, and never a nested box inside a box.

Use the existing text-hierarchy tokens for this rather than inventing new ones: --color-text-primary for the one thing that should read as bright (a name), --color-text-secondary / --color-text-tertiary for everything that should recede, --color-border for hairlines. These are already contrast-checked; a new ad hoc muted colour is not.

Reserve the accent colour for state, not decoration. Terminal green means this is active/healthy right now, and only that. It is not a border, a background tint, or a default text colour applied everywhere on principle — if it is everywhere, it stops signalling anything, which is exactly why the fake per-row traffic lights were a defect and not a style choice.

At most one boxed element per view. If a page has two or more .terminal-window panels, ask which one actually deserves it before adding a third.

Addendum: cards vs. tables (added after dashboard/workers shipped)

The first pass of this ADR converted dashboard/workers to a hairline table, correct by the letter of the rule above and wrong for the content. Chad’s correction: “workers needs to be much more interesting. this is where we would use cards.”

The rule was never “tables everywhere” – it’s that chrome must earn its place, and the right shape follows what the content actually is, not just its item count:

  • Table: many similar items you scan and compare against each other – 42 projects, a list of documents, a row of requests. Hierarchy from type and space, exactly as written above.
  • Cards: a handful of rich, individual things where you want to see inside each one – a few worker machines, each with its own live activity, resource use, and assigned work. Forcing these into table rows discards the one thing that made them worth looking at.

A card is not a reversion to the original defect. The thing ADR-003 killed was decorative chrome repeated per row with no informational content – a border, a title bar, three dots that meant nothing, 42 times. A card that holds real, current information (what a machine is doing right now, not a status pill) earns its border by showing something; it isn’t punctuation pretending to be paper. See dashboard/workers/index.astro for the applied version – real activity text pulled from live session data, a distinct color for a session blocked on a human, motion gated on genuine busy state, and unreported metrics shown as unavailable rather than faked as zero.

Reference implementation

/styleguide/terminal-chrome renders the current per-card treatment next to a table built to this rule, using the same project data shape, at desktop and mobile. Read it alongside this entry rather than reimplementing the rule from prose alone — the projects table itself is being rebuilt separately to this reference, not by editing that page.

Not done in this entry

This ADR documents the rule and demonstrates it; it does not rewrite every file listed above. The projects list is being rebuilt to this standard as its own piece of work. The remaining single-panel and stacked-panel cases (dashboard/index.astro‘s stat tiles) are not fixed here and should be treated as a backlog against this rule, not a silently accepted exception to it. Settings pages, dashboard/workers, and dashboard/documents have since been addressed (see the addendum above for workers’ cards-not-table treatment) and are no longer part of this backlog.