Build a dated model card
<p>Preserve original inputs and outputs, record dates and operators, and separate facts, observations, creative preferences, and unknowns. Do not infer access, price, quality, rights, or repeatability from circulation.</p><p>Keep the record useful to a second reviewer: identify the exact phrase being investigated, state what would count as a successful observation, and distinguish a missing artifact from a negative result. Note the account, region, software version, input permissions, and delivery context when they affect interpretation. If a source changes, keep the earlier observation and add a new dated row rather than silently rewriting history. This makes the worksheet portable across a direct test, an editorial comparison, or a conventional production fallback without implying that any unverified service is available through SEELE. Add a clear owner for each unresolved question, the next source or observation to obtain, and the date by which a stale decision should be reopened. For visual work, retain the original files and export settings; for audio, retain timing, channel, and consent notes; for software or model work, retain the exact interface label and dependency versions. Record failed attempts without treating them as proof of universal failure. When a result is useful only for planning, label it as planning evidence. When a result is suitable for publication, separately review rights, disclosure, factual claims, and destination requirements before delivery. Add a verification table that names the exact query, the source that introduced it, the first-party page or direct test still needed, the reviewer responsible for that check, and the date when the observation becomes stale. Treat a provider name, a model name, a wrapper name, and a feature name as separate identity fields; matching words are not proof that they refer to the same system. Record whether the reader can reproduce the observation with an authorized account and whether the result depends on a hidden preset, an unavailable asset, or an undocumented post-processing step. Compare failures as carefully as successes, but keep their scope local to the recorded attempt. Before a handoff, have a second reviewer challenge the strongest inference, locate every unsupported adjective, and confirm that the fallback remains useful without the named service. This produces an auditable model-review packet while preserving uncertainty instead of manufacturing a capability claim.</p><ul><li>Freeze the brief and source materials before testing.</li><li>Record every material transform and the evidence it supports.</li><li>Keep a fallback path when the named product or behavior cannot be verified.</li></ul><p>For a model review, also capture the exact provider label, endpoint or interface, version date, account or region, hardware path, prompt and input files, output settings, latency, failure mode, and authorization status. Compare only matched attempts and label observations as unresolved when the named model cannot be directly identified. Keep a separate access log for sign-in, quota, endpoint response, and billing or entitlement observations. Record whether a result came from a hosted interface, an API, a local runtime, or a third-party wrapper. Do not merge those paths into one model identity. A useful review ends with a reproducible handoff: another reviewer should know which evidence is confirmed, which observation is provisional, which test was not run, and what must be checked again before publication. Include a compact comparison table for the exact interface label, version, input class, output dimensions, timing, failure mode, and evidence status. Keep provider marketing language separate from observed behavior, and record the smallest next test that could change the conclusion. If the named model cannot be identified, preserve that uncertainty rather than substituting a nearby model or wrapper. This protects readers from false equivalence and keeps the route useful as a research worksheet.</p>