THE MIRROR / 05 / HOW IT WORKS

A mirror, with the wiring exposed.

The web reveals more than a page. THE MIRROR makes ordinary browser and network observations visible, and gives you a way to inspect their limits.

Four ways a website can know something

A · Automatically exposed. Your request arrives with headers. The hosting edge can see your connection address and may provide a coarse location, network organization and protocol details. Arbitrary forwarded headers are never treated as trustworthy.

B · Script-readable. JavaScript can read your viewport, language list, timezone and supported APIs. These observations stay in the current tab unless you deliberately save a reflection. Capabilities vary between browsers.

C · Permission-protected. Location, camera, microphone and other sensitive APIs require a deliberate action, and usually a browser prompt. Each test in Permission Lab stands on its own. Media and precise coordinates are never uploaded or included in a saved reflection.

D · Inferred. Browser names, rendering engines, IP location and similarity scores are interpretations. They may be wrong. Every signal’s detail view explains its source and destination.

What an interaction means

The live instrument counts pointer moves at most 20 times per second. Keyboard listeners increment a counter; they never inspect which key was pressed. Form listeners record only field type, focus duration and whether something changed. They do not read field contents.

Visible time and hidden time come from the Page Visibility API. A session becomes idle after 30 seconds without activity or while the page is hidden. Browsers may throttle timers or suspend tabs; elapsed totals are corrected when the page resumes. The event stream keeps at most 100 entries in local memory.

A reflection is an explicit snapshot

Use ChatGPT sign-in to identify your account. Saving creates an application profile and associates only the approved environment values, permission states, route durations and aggregate counts with it. There is no hidden retroactive association across anonymous visits.

Manual saves capture one snapshot. Automatic saving is an explicit account preference; when enabled it updates a snapshot every 30 seconds while observation is active. Each origin has independent browser state and a separately scoped machine grouping key.

Resemblance is not identity

Comparison gives browser a weight of 2, and platform, language list, timezone, screen dimensions, pixel ratio, graphics renderer, viewport and public IP a weight of 1 each. An exact match earns its weight. Missing or withheld pairs are excluded from the denominator.

Two different people may share every measured value. One person can change many values by moving a window or using a VPN. The optional similarity experiment compares only your own saved history and makes no claim of uniqueness or cross-site identity.

How machines are classified

A known crawler signature is a claim, not verification. A declared headless browser receives a heuristic confidence of 85%; a scripted HTTP library, 80%; a known crawler signature, 70%; an unknown bot-like name, 60%. These are transparent rule scores, not calibrated probabilities.

A trusted Cloudflare verified-bot signal earns a 95% rule score for automation. Even then, the claimed name is not independently confirmed. DNS verification and published-IP-range verification are not enabled in this deployment. Anonymous browser-looking requests remain “likely human” with low confidence; absence of a JavaScript beacon never proves automation.

Machine grouping uses a secret-keyed hash of hostname, connection address, user agent and a 30-minute time window. The database keeps the resulting specimen ID, not the raw address. Shared networks may merge clients and time-window boundaries may split a crawler.

The labyrinth ends

The default graph contains 256 deterministic, clearly synthetic entities with at most four outward links per record and a maximum structural depth of four. IDs outside the configured range return 404. Representations are bounded HTML, JSON, JSON-LD, XML and RSS; unknown formats are rejected. There are no infinite responses, unbounded pagination, redirect loops, decompression tricks or deliberately expensive computations.

The experiment permits at most 60 machine-route requests per connection per minute. Logs have a daily per-category cap of 20,000 requests. A repeated route in a specimen’s recent 25-request window is displayed as a traversal loop. It is evidence of revisiting, not evidence of an agent getting stuck.

Robots policy is not a lock

robots.txt disallows API routes, the admin route, and a harmless synthetic test room. The room contains no sensitive data. The private dashboard reports whether the same heuristic specimen fetched the policy before requesting the disallowed room; missing evidence is explicitly labeled.

Browser differences are part of the exhibit

Safari and Firefox intentionally expose fewer hardware and connection details than some Chromium browsers. Mobile sensors may need explicit permission. Media devices can be unavailable in embedded browsers, insecure contexts, private browsing or operating-system settings. “Not exposed” is a useful result, not an error to work around.

Browser permission reference · Hosting request metadata · Inspect and delete your data