Skip to main content
ControlTowerAI

WHAT’S BEHIND IT

How It Works

Four short answers: why we observe AI behavior, what we measure, how the measurement is produced, and what it does not say.

01 · PURPOSE

Why observe AI systems?

Artificial intelligence systems are probabilistic, evolving and complex. Unlike fully fixed software, they can produce different answers under comparable conditions, and their behaviour can change over time.

For four months, the NeoMundi Observatory has been running a continuous programme observing AI systems. AI Weather™ is its public demonstrator: every day, it makes visible what these systems do, how they vary and how their behaviour shifts.

Its message is simple: to judge and govern an AI, you first have to be able to measure it. Discovering a drift after an incident is like realising you needed an umbrella once you are already standing in the rain.

NeoMundi provides the measurement layer missing between AI and the systems that use it. During execution, it produces a traceable quality and risk signal. That signal can then be audited, compared and governed.

AI Weather™ shows the principle publicly. NeoMundi’s infrastructure applies it to real-world use: measuring during execution rather than hunting for the incident afterwards.

Read more: Why measure AI?

Four levels to guide attention

“An AI can be wrong”: that warning applies just as much to individuals as to the companies deploying these systems. Without measurement during execution, controlling the risk means relying on intuition alone.

Each colour indicates an observed behavioural condition. It shows where to look, without deciding on the reader’s behalf.

Normal

FR: Normal

Today’s observations show no concerning change. Keep checking answers that matter.

Watch

FR: Vigilance

A first change is appearing. For instance, the model answers the same question less consistently. Compare its answers before relying on them.

Warning

FR: Alerte

The change is more pronounced or it recurs. An answer that previously looked reliable may become less predictable: check it against another source.

Critical

FR: Critique

A significant disruption was observed. Avoid relying on this result for an important decision without human analysis.

Green does not mean “no risk” or “safe model”: it only describes the observations covered by the protocol. These levels support human judgment and are neither a certification nor a guarantee.

02 · WHAT WE MEASURE

What we measure

Here is how an observation becomes a coloured state, and what that state actually covers. Nothing here is a performance or quality score.

The measurement in four steps

1. The same questions

Every day, the same sentinel questions are put to the systems in the panel, under comparable conditions.

2. Measurement and trace

Each answer received is measured and timestamped in UTC, and the corresponding trace is kept.

3. Comparison over time

The day’s measurements are compared with those of previous days, which is what makes variations visible.

4. A coloured state, coverage visible

The result is published as a coloured state, together with the share of observations actually carried out. Incomplete coverage stays explicitly visible.

What the NeoMundi layer can measure

Depending on the systems and the access available, the NeoMundi measurement layer can produce signals for stability, semantic consistency and factuality, as well as token, cost and latency measurements. These are general capabilities of the layer, available for a business integration. They are not all switched on for the public station.

What the public station measures today

As of today, AI Weather™ uses only the three notions below, plus panel coverage. No other metric is published on this station.

Stability

How consistently a system answers across repeated runs of the same sentinel protocol.

Variation

The measurable spread between several answers obtained under controlled, comparable conditions.

Behavioral drift

A lasting change in observed behaviour across several successive measurements, distinct from a one-off variation.

03 · HOW IT’S PRODUCED

How the measurement is produced

WEATHER-SENTINEL, the fixed protocol behind every card on the wall.

01

Fixed panel, fixed questions

The same panel of systems receives the same sentinel questions every day, under documented parameters and in conditions kept as comparable as possible.

02

Repeated runs

Each system is subject to 30 runs per day: a main sentinel question repeated 23 times, and a fixed complementary question repeated 7 times.

03

Comparison against the reference

The day’s observations are compared against the system’s recent history. Stability, the size of the variation, its persistence and the available coverage all contribute to determining the published condition.

04

Verifiable publication

The result is written to a timestamped capsule, chained by cryptographic hash, then published as JSON data and reusable widgets.

04 · LIMITS

What the measurement does not say

AI Weather™ observes and measures behavioural signals. It does not turn measurement into a verdict.

  • It does not rank systems and does not name a “best” model.
  • It neither certifies nor guarantees their future behaviour.
  • It does not judge whether an individual answer is true or false.
  • It replaces neither audit, nor compliance, nor governance, nor human judgment.
  • An observed level is a signal to pay attention, never an automatic decision.

NeoMundi provides the signal. The reader interprets it and decides.

Why this is not a benchmark

A benchmark usually evaluates a model’s performance on a set of tasks at a single point in time.

NeoMundi measures something else: the continuity of a system’s behaviour over time and while it runs.

A model can keep good average performance while becoming less stable or shifting regime. It is that invisible evolution that a longitudinal measurement makes visible.

A NEW MEASUREMENT LAYER

Measure before governing

AI Weather™ is the public demonstrator of a simple idea: you cannot durably govern what you do not measure.

NeoMundi is building the measurement layer missing between AI systems and the organisations that use them. An independent primitive turns execution observations into a traceable, comparable and auditable signal.

This approach rests on four differences:

  • Measuring during execution, rather than discovering deviations after an incident.
  • Building a longitudinal memory, so that changes in behaviour over time become visible.
  • Connecting a single measurement to several uses and infrastructures, without rebuilding the system for every integration.
  • Producing a machine-readable trace, stating what was measured, when, in what context and under which version.

An infrastructure built in the open

The NeoMundi Observatory runs a continuous programme of observation and of measurement-based critical-thinking outreach.

The Observatory is supported by Infomaniak. NeoMundi technology is supported by NVIDIA Inception et l’OVHcloud Startup Program.

AI Weather™ does not ask anyone to trust a ranking. It gives everyone a documented starting point from which to observe, compare, question and decide with more discernment.

Support and programmes

  • Logo d'Infomaniak
  • Logo de NVIDIA Inception
  • Logo d'OVHcloud Startup Program

Need a specific format?

We can look into a format adapted to your media outlet, your institution or your digital environment.

Press & media

The How It Works page brings together the method, the context and reusable resources for journalists, researchers and the simply curious.