Earth or Mars
A Head-to-Head Test of the 3x3 Framework Against Frontier AI on Five Planetary Surface Images
Visual scenes are a demanding test of any framework that claims to model systems rather than lists of objects. An image has parts, behaviors, and — most importantly — emergent properties that only appear when the parts interact. That is exactly what the 3x3 framework was built to handle.
This is a worked example. The framework is general: it models markets, supply chains, technologies, organizations, and regions. The image domain gives us somewhere to test it against ground truth and against frontier AI systems reasoning natively over the same pictures.
For the methodology in plain English, see The 3x3 Methodology →.
The test
Five planetary surface images. Each requires a structural classification — Earth or Mars — that depends on context rather than on object labels. Three approaches were scored against ground truth.
| Approach | Score |
|---|---|
| 3x3 framework | 5 / 5 |
| ChatGPT 5.2 | 4 / 5 |
| Claude Sonnet 4.5 | 3 / 5 |
The score is not the point. The error mode is.
Both frontier models rely on object-level features — texture, color, surface appearance — and miss the contextual cues (atmosphere, scale, weathering, biological or human traces) that distinguish the two regimes. The 3x3 framework’s structural prior makes those contextual cues visible.
Why the comparison matters
Image classification is the easiest place to compare a structural approach to a pattern-matching one, because ground truth is unambiguous and the failure modes are visible.
The frontier models in the comparison are remarkable systems. They reason. They generalize. They also miss things, and the things they miss are the situational and contextual cues that the 3x3 framework is built to make visible.
The pattern is consistent across both models. When the classification depends on a list of object features — texture, color, surface appearance — the models do well. When it depends on the interaction of features — atmosphere plus shadow, scale plus weathering, mechanism plus context — they miss. The 3x3 framework’s structural approach, entities and behaviors and the emergent properties arising from their interaction, is what makes that interaction visible.
This is the same advantage that produces a 30-to-90-day lead time on a regime change in a market, a supply chain, or a defense theater. The framework does not see objects. It sees the situation the objects are in. In an image, the situation is the planetary context. In a market, it is the structural state the market is sitting in. The methodology does not change; the input changes.
The error mode, in detail
The most instructive case is image 2 — a Saharan dune under a blue sky. The blue sky is a contextual cue. The dune texture is an object feature. The two combine into a clear Earth classification. A pattern-matching model, looking at the dune alone, sees a Mars-like signature and stops. The 3x3 framework sees the dune, sees the sky, and sees that the interaction of the two is decisive.
Claude Sonnet 4.5 misclassified images 1 and 4 — both cases where contextual cues (rover hardware, biological desiccation patterns) outweighed the object features. ChatGPT 5.2 misclassified image 2. Both models relied on object-level reasoning. The 3x3 framework’s structural prior avoided the trap on every image.
How the 3x3 model helps
We structure the evidence into three layers:
- Entities (O): observable objects and surfaces — regolith, rocks, tracks, shadows
- Behaviors (B): the physical processes shaping them — erosion, sediment flow, compaction, illumination
- Emergents (e): higher-level cues that appear from interactions — patterning, depth, directionality
When an image is dominated by mechanism-only evidence and lacks biological or human traces, the model shifts toward a Mars classification. When we see water-weathered surfaces, vegetation patterns, or human scale and context, it shifts toward Earth.
Decision cues used in practice
Typical signals that push the classification toward Mars:
- Dry, fine-grained regolith with uniform dust tones
- Sparse, angular rocks without water-weathering
- Strong, sharp shadows in a thin atmosphere
- Rover-like tracks or hardware shadows without surrounding context
Typical signals that push toward Earth:
- Vegetation or biological texture
- Mixed mineral colors with moisture cues
- Weathering from water or wind — rounded stones, sediment layers
- Human artifacts or scale references — roads, fences, footprints
The image set
All five correctly identified by the framework. Select any image for its full O/B/e analysis.
| Image | Verdict | Primary cue | 3x3 | ChatGPT 5.2 | Claude Sonnet 4.5 |
|---|---|---|---|---|---|
| Mars | Regolith + rover-like tracks and shadow — view analysis | ✓ | ✓ | ✗ | |
| Earth | Blue sky + terrestrial dune texture — view analysis | ✓ | ✗ | ✓ | |
| Mars | Rover-like shadow + dry regolith — view analysis | ✓ | ✓ | ✓ | |
| Earth | Mudflat desiccation polygons — view analysis | ✓ | ✓ | ✗ | |
| Mars | Large dune scale + dry aeolian signature — view analysis | ✓ | ✓ | ✓ |
What this is, and is not
This is a worked example of the methodology in a domain where the answers are checkable. It is not a product page. The 3x3 Institute does not sell image analysis — we offer a methodology for modeling complex systems and forecasting their trajectory.
If the image work is what brought you here: the same framework that scored 5 of 5 on a planetary classification is the framework that models a market in transition, a supply chain at risk, and a patent portfolio’s vulnerabilities. The image domain is where the methodology is most testable. The strategic domains are where it is most useful.
To see the same framework applied to strategic questions, see Applied Analyses →. To start a conversation about your domain, get in touch →.