Use Case

Human-AI Interaction Research

Capture, code and analyse how people actually use AI in professional workflows — with the rigour to generate evidence that stands up to scrutiny.

The use case

Capturing, coding and analysing the cognitive and behavioural dynamics of AI-mediated work at research rigour.

Problem

The gap between use and good use

Organisations deploying AI tools rarely know how their people are actually using them — which tasks they're offloading, how they're prompting, whether outputs are being evaluated critically or accepted uncritically. General analytics don't answer these questions.

Workflow

Measure, intervene, measure again

Capture AI interaction sessions in Captures during normal document work. Define coding schemes for cognitive acts, offloading behaviours and reasoning patterns in Library. Apply the framework to the corpus in Observatory. Intervene. Repeat.

Output

Evidence that does two jobs

Pre/post corpora with preserved inference records support both demonstration of intervention effectiveness and validity evidence for the measurement framework itself. The same study generates client-facing results and research-grade output.

Human Behaviour AI

A partnership for measurement-grounded AI implementation consulting.

Human Behaviour AI is a consultancy founded by Fendi and Jason, built on a research tradition of cognitive offloading and augmentation theory. Their work addresses a specific gap: organisations investing in AI tools don't know how their people are actually using them. HBAI's approach is to measure first — capturing real AI interaction data in professional workflows — and then design interventions based on observed patterns rather than assumptions about what good AI use looks like.

Their framework operationalises contemporary cognitive science: offloading frequency, prompt strategy sophistication, output evaluation depth, and revision behaviour are treated not as soft observations but as measurable dimensions with theoretical grounding in the distributed cognition and extended mind traditions. What HBAI needed was infrastructure that could handle that measurement at the scale and rigour their framework demands.

The study

A pre/post intervention design in a professional services firm — measuring how AI is used in document workflows before and after an HBAI implementation programme.

The client was a mid-size professional services firm that had rolled out AI assistants across its document-intensive workflows — research synthesis, briefing documents, client communication — but had no empirical basis for evaluating whether their internal AI training was producing meaningful change in how people worked, or whether it was changing the right things. HBAI was brought in to run an implementation programme and generate the evidence to know whether it worked.

Interface provided the measurement infrastructure. Across a four-week pre-period, 24 participants across three practice areas worked with AI in their normal document workflows, with sessions captured through Interface Captures. HBAI's assessment framework — loaded into Library as a structured set of coding dimensions — was applied to the resulting corpus through Observatory, producing a baseline profile across four dimensions: offloading frequency, prompt strategy, output evaluation depth, and revision behaviour.

After a three-day HBAI workshop programme, the same capture-and-code methodology ran for a further four weeks. All inference records were preserved with prompt version, model, timestamp and coder — making the two corpora directly comparable and the analysis auditable.

Phase Duration Method Output
Pre-measurement 4 weeks AI interaction capture in normal document workflows Baseline cognitive workflow profile across 24 participants
Intervention 3 days HBAI implementation programme Targeted shifts in prompting strategy and output evaluation behaviour
Post-measurement 4 weeks Same capture-and-code methodology, same framework version Post-intervention profile, directly comparable to baseline
Analysis Pre/post comparison across framework dimensions; discriminant validity check Intervention evidence + framework validity evidence

Results

The study generated two distinct bodies of evidence from one corpus.

Intervention evidence

Meaningful shifts on targeted dimensions

Post-workshop, participants showed a 34% increase in instances of explicit critical review of AI outputs, and a statistically significant improvement in prompt strategy sophistication across all three practice areas (effect size d = 0.72). Session length and general writing quality — control dimensions the workshop wasn't designed to affect — showed no significant change.

Validity evidence

Framework sensitivity confirmed

The discriminant pattern matters as much as the changes themselves: HBAI's framework was sensitive to exactly the dimensions the intervention targeted, and insensitive to the ones it didn't. That specificity is what distinguishes a valid measurement instrument from one that moves with everything and therefore explains nothing.

Infrastructure

Auditable from first capture to final export

Every inference in the coded corpus carries the framework version, prompt version, model, timestamp and coder. The pre and post corpora can be compared with confidence because the measurement conditions are documented and held constant. The analysis is exportable for publication or client reporting.

"We've always believed the gap between using AI and using AI well is measurable. Interface gave us the infrastructure to measure it with the rigour our framework demands. The pre/post corpus generated evidence we can stand behind — both for our clients and for the validity of our own approach."

Jason · Co-founder, Human Behaviour AI

Principles in play

The design principles most active in this use case.

Principle

Augmentation

The study is about augmentation — measuring whether and how AI is genuinely extending human capability in professional work rather than short-circuiting it. Interface is the tool studying the phenomenon it embodies.

Open →
Principle

Confidence

Validity evidence is confidence evidence. The discriminant pattern — the framework moves where it should move and doesn't where it shouldn't — is what makes confidence in the intervention results warranted.

Open →
Principle

Transparency

Every inference in both corpora is traceable: which framework version, which prompt, which model, which coder, when. The transparency of the measurement process is what makes the results defensible.

Open →

Research foundations

The academic traditions this work sits within.

Research

Human-AI Interaction

From Licklider's symbiosis to contemporary trust calibration research — the tradition of studying human and machine as a coupled cognitive system.

Open →
Research

Distributed Cognition

Hutchins and Kirsh on epistemic actions and cognitive offloading — the theoretical basis for understanding what it means to distribute work between human and AI.

Open →
Research

Validation Theory

Kane's argument-based approach to validation: the discriminant evidence from this study is a direct application of construct validity methodology to a consultancy measurement framework.

Open →