Protocol design

AI Applications Laboratory

Problem

AI applications may improve information-processing efficiency, but opaque sources, automated error amplification, bias, privacy leakage, and responsibility transfer can undermine value.

Research question

Can an AI workflow with provenance, task boundaries, and human review improve accuracy and efficiency on specified tasks without increasing severe errors?

Testable hypothesis

The following is a testable proposition, not an established result.

Compared with the task baseline, a constrained AI workflow will improve preregistered accuracy or time measures while keeping severe errors below the safety threshold; one assessment without net improvement is not decisive, and preregistered thresholds plus repeated validation determine whether evidence supports progression, while a safety-threshold breach triggers stopping or revision.

Measures

  • Task accuracy, completion time, calibration, and verifiable-source rate
  • Severe errors, human corrections, refusals, and escalation handling
  • Bias, privacy, security, and usability across user groups

Method

First define the task, baseline, error tiers, and stopping rules, then run offline benchmarks with synthetic or authorized data; high-impact settings require independent safety review before any constrained pilot.

Current evidence

The following states the current record and evidence types separately from the hypothesis.

The current record is an application-evaluation framework and protocol design, with no claim of a deployed product, real-user data, or performance result.

  • Theory source
  • Testable hypothesis
  • Research protocol

Evidence gaps

  • Defined use cases, representative benchmarks, error costs, and acceptance thresholds are still missing
  • Bias, privacy, security, human-factors, and continuous-monitoring validation are still missing

Milestones

  1. Complete the first use cases, baselines, measures, and stopping rules
  2. Complete offline benchmarks, red-team testing, and independent safety review
  3. Decide on pilots only for low-risk settings that pass review

Governance

For each use case, specify accountability, permitted data, model version, and human-review points; reassess after model or prompt changes and do not transfer responsibility to automation.

Ethics

Follow data minimization, informed use, fairness, and appeal principles; unvalidated AI must not autonomously make high-impact medical, legal, financial, or personal-safety decisions.

Partner needs

  • Task-domain, AI-evaluation, human-factors, and independent-audit institutions
  • Privacy, security, accessibility, and affected-user representatives

Planned public outputs

  • Use-case cards, evaluation datasheets, and safety-threshold registry
  • Offline evaluation, red-team testing, and scope-of-use report

Last reviewed