How it works

How it works

Written for someone who has never seen this system. No background assumed.

What this is, in one paragraph

Cameras record people doing everyday tasks. This system watches that footage and answers three questions: who was in the room, how they seemed, and whether they completed a task properly. It is built for supported-living settings, where a support worker cannot stand over every client all day but still needs to know whether the vegetables got washed before they got cooked.

The one thing to take away
The system never guesses whether a task was done correctly. It identifies what it sees, then checks that against rules you wrote down in advance. Those are two different jobs and keeping them apart is what makes the answer defensible.

The model decides one thing; code decides the rest

The obvious approach is to show an AI model a video and ask “was this done correctly?” That fails in a specific and damaging way: you get a different answer on different days, and when it says no, nobody can explain why to the person it is about.

So the work is split in two. The model is asked one narrow question about each few seconds of video: which step of the task is this? That is a perception question, which is what models are good at. Deciding whether the resulting order is acceptable happens in ordinary code, against rules written down before the footage existed.

The model decides
  • Which step a window of video shows
  • How confident it is about that
  • A one-line reason you can read
Code decides
  • Whether a required step never happened
  • Whether the order broke a rule
  • Whether a person needs to look at it

The practical benefit: if the verdict is wrong, you can point at the exact five seconds the model misread. If the rules are wrong, you can change them in a form. Neither failure is a mystery.

A workflow is steps plus an order

A workflow describes one task. Preparing vegetables has four steps: take, wash, cut, cook. But listing the steps is not enough, because the order matters and it is not a single fixed sequence. Washing and cutting can happen either way round. Cooking must come last. Taking must come first.

Rather than writing out every acceptable sequence, you state the constraints: take before everything, cook after everything. Every valid order falls out of those rules, and every invalid one is caught.

✓take → wash → cut → cookeverything in a sensible order
✓take → cut → wash → cookwashing after cutting is still fine
✕take → cook → wash → cutcooked before washing or cutting
✕take → cook → cut → washcooked first, which is the whole problem

You build this on the Workflows page by ticking a grid. No code, and you can try a sequence against your rules before saving to check they say what you meant.

Your first run, step by step

If you have just opened this system and want to see it do something, this is the shortest path.

  1. 1
    Bring in some footage
    Go to Sources. Paste a YouTube link of someone doing a task, or point it at a camera, or upload a file. The engine samples it into still frames and builds a contact sheet.
  2. 2
    Check the sampling looks right
    Open the clip and look at the contact sheet. If the important moments are missing, ingest again with more frames. This takes ten seconds and saves a confusing result later.
  3. 3
    Define what the task should look like
    On Workflows, list the steps and tick which must come before which. Write a clear hint for each step: that hint is literally what the model is told to look for.
  4. 4
    Read the verdict
    The clip page shows what the model assigned to every window, how sure it was, and why. Then the verdict, and if it failed, exactly which rule broke and at what timestamp.
  5. 5
    Correct it if it is wrong
    Anything uncertain lands in Review. Accept, correct or reject it. Your decision is recorded against your name, which is what makes the record auditable.

Confidence, and why some runs get flagged

Every label the model produces carries a confidence figure. High confidence means the frames clearly showed that step. Low confidence usually means the window caught a transition, or the camera angle hid the hands, or the step genuinely was ambiguous.

A run is flagged for human review when any step fell below 60 percent, or when the verdict was a failure. A failure is flagged even at high confidence because a failed verdict is the kind that has consequences for a person, and those should not go unread.

Windows the model calls idle
Not every second of footage is part of the task. Pauses, walking away and tidying up are marked idle and dropped. They carry no ordering information, and keeping them would make a coffee break look like a missing step.

Using your own database, storage and cameras

By default the engine keeps everything on its own disk. That is fine while you are trying it and wrong for anything real, because a redeploy wipes it.

On the Connections page you can register a PostgreSQL database (Neon, Cloud SQL, Supabase or self-hosted), a Google Cloud Storage bucket for media, and any number of camera streams. Each one is tested against the real service before it can be saved, so a mistyped password fails in front of you rather than three days later.

Secrets are stored write-only. Once saved, the console shows them masked and there is no way to read the original back out, including for you.

Where client video goes

Nowhere, on a properly configured site installation. The engine can be started in on-premise mode, and in that mode it refuses to start at all if it is also configured to use a hosted AI model. That is a deliberate hard failure: a misconfigured site should not quietly start sending footage of vulnerable people abroad.

What may cross to a hosted service is anonymised text: step names, timestamps, verdicts, counts. No names, no faces, no images.

About the sample clips in this console
The reference clips here are public videos used during development, which is why a hosted model was allowed to read them. They are not client footage and the rule above does not apply to them.

What this will not do

Being clear about the edges matters more than sounding capable.

  • It does not judge quality. It can tell you the vegetables were cut; it cannot tell you they were cut well.
  • It does not diagnose anything. A distress score is an observation to bring to a person, not a clinical finding.
  • It does not watch continuously and decide alone. It processes clips, and a person confirms anything consequential.
  • It is only as good as the workflow you wrote. Vague step hints produce vague labels.

Glossary