
“The air on top doesn’t need to catch up at all,” wrote Claude Fable 5—adding that it “usually gets to the back first.” That is a remarkably direct correction of one of the most familiar bad explanations of flight: the idea that air split by a wing must reunite at its trailing edge at the same time. In this one response, Fable paired the correction with a child-friendly central idea: wings push air downward, and air pushes the wing upward.
Disclosure: I work with OrcaRouter and used it to run this evaluation. One OpenAI-compatible key gave me access to every model in this test, without changing how any model answered.
AI-generated illustration featuring official model logos; logos and model names are used descriptively and remain the property of their respective owners.
Compare the models in this article through OrcaRouter’s model catalog.
That was the narrow challenge in this case study. Eight models received exactly one prompt: “Explain to a 10-year-old why airplanes can fly. Keep it accurate — no myths.” The evidence is therefore one child-facing science explanation per model, not a general measure of scientific ability or a stable ranking of these systems. The prompt also specifically tested whether a model would avoid the equal-transit-time misconception.
The eight selected gateway/API calls were made in run 20260717_120445, with search requested and effective set to False for each. They covered GPT-5.6 Sol, Terra and Luna; Claude Opus 4.8 and Fable 5; Grok 4.5; Gemini 3.5 Flash; and GLM-5.2. All eight calls returned ok.
The common answer: wings turn air downward
The striking result was convergence, not a clean winner. Every response built its explanation around lift, forward motion and a wing’s interaction with air. Most plainly said that the wing pushes air downward and receives an upward push in return.
GPT-5.6 Terra opened with precisely that mechanism: “Airplanes fly because their wings push air downward.” It then explained that the wing’s tilt, or angle of attack, redirects air behind it downward. GPT-5.6 Luna presented the same sequence as four steps, beginning with engines moving the aircraft forward and ending with air pushing upward on the wing.
Claude Opus 4.8 and GLM-5.2 used a familiar physical demonstration: tilt a hand into the wind from a moving car window and it rises. Gemini 3.5 Flash organized the answer around two tug-of-war pairs—thrust versus drag and lift versus gravity—before explaining that a tilted wing pushes air down.
That shared structure makes these answers useful as a use-case profile: for a basic, child-directed explanation, all eight observed replies gave a reader a workable core mental model rather than relying solely on wing curvature.
Who addressed the myth—and how?
Four responses explicitly named or corrected the equal-transit-time story.
- GPT-5.6 Sol said streams over and under the wing “do not have to meet again at the same time.”
- GPT-5.6 Luna likewise said lift is not caused by particles having to meet at the back simultaneously.
- Claude Fable 5 directly rejected the “catch up” claim and said upper airflow usually reaches the back first.
- Grok 4.5 called the equal-transit-time idea untrue and said it was skipping it.
The other four did not repeat that misconception in their answers. Claude Opus 4.8 instead cautioned that saying planes fly only because air moves faster over the top is incomplete and often explained incorrectly. GPT-5.6 Terra used the fact that airplanes can fly upside down for a while to emphasize the importance of wing angle and deflected air, rather than just a curved upper surface.
A compact view of the observed patterns
| Model | Main explanatory pattern in this one reply | Explicit equal-transit-time correction? |
|---|---|---|
| GPT-5.6 Sol | Downward-deflected air, pressure, four forces | Yes |
| GPT-5.6 Terra | Downward-deflected air, angle of attack, upside-down-flight clue | No |
| GPT-5.6 Luna | Step-by-step downward-deflected-air explanation | Yes |
| Claude Fable 5 | Hand-out-the-window analogy, downwash and pressure | Yes |
| Claude Opus 4.8 | Hand-out-the-window analogy and downwash | Indirectly addresses a related oversimplification |
| Grok 4.5 | Long-form two-part account: downwash and pressure | Yes |
| Gemini 3.5 Flash | Four forces and a two-tug-of-war structure | No |
| GLM-5.2 | Four forces, hand analogy, downwash and pressure | No |
There are real stylistic trade-offs here. Fable and Opus used concrete hand-in-the-wind imagery that may be easy for a child to picture. Luna used numbered steps and stayed comparatively compact. Grok and Gemini gave more extended tours through the forces acting on a plane. Those are observations about these individual outputs, not proof that one model will always be clearer, shorter or more accurate.

Reproducible data figure from this article’s selected API records; it is not a general science-capability ranking.
What the scoring did—and did not—say
The selected successful row for every model received an accuracy score of 5 under the v3 judge prompt. But this was a reference-guided DeepSeek v3 judgment, not expert physics education review. It should be read as an audit signal that the eight answers matched the evaluation’s reference-guided standard, not as certification that each explanation is ideal for every 10-year-old.
This distinction matters. A model can include good physical ideas while still choosing a level of detail, analogy or vocabulary that one child finds confusing. Several answers introduced terms such as “angle of attack,” “Bernoulli’s principle,” or “Newton’s third law.” Whether that is helpful depends on the reader and the teaching setting; this test did not measure comprehension with children.
The study also should not be folded into broad AI benchmark claims. Humanity’s Last Exam is a closed-ended academic benchmark and is context only; it does not validate a child-facing explanation of aerodynamics.
Runtime observations are not a leaderboard
In these selected calls, Gemini 3.5 Flash returned in 7.53 seconds, while GPT-5.6 Terra took 68.67 seconds. Those are single-call latency observations, not evidence that one model is generally faster. Likewise, the run’s recorded billed amounts and prompt-token counts varied, but this one-question comparison cannot establish a general cost or efficiency winner.
The model labels also must not be read as equal effort or compute settings across vendors. And because these were gateway/API observations, they do not describe how consumer subscription products behave.
Practical takeaway
For a parent, teacher or product designer who needs a quick airplane explanation, the best practical check is simple: look for an answer that says wings redirect air downward, connects that to an upward force on the wing, and does not claim that split air must reunite at the back of the wing at the same time.
In this small test, every observed answer cleared that basic bar, while GPT-5.6 Sol, GPT-5.6 Luna, Claude Fable 5 and Grok 4.5 made the myth correction explicit. If clarity for a particular child matters most, a short step-by-step answer or a hand-in-the-wind analogy may be more useful than a longer technical tour.

Editorial illustration; it frames a child-facing explanation of flight and is not test evidence.
Limitations
- This is one prompt and one selected child-facing science explanation per model.
- The prompt was designed to test avoidance of the equal-transit-time misconception, so it cannot represent general science or teaching performance.
- Each question-model cell has a sample size of one.
- The v3 reference-guided DeepSeek judgments are not expert physics education review.
- Observed API latency, tokens, billing and reliability are descriptive only, and cannot be generalized to consumer products or broader model behavior.
Explore the Models
Explore the current catalog on OrcaRouter Models.
This evaluation was run through OrcaRouter. The author works with OrcaRouter; model access does not imply affiliation with, endorsement by, or sponsorship from model providers.
Model names and logos are used descriptively. All trademarks belong to their respective owners.
Sources
- Humanity’s Last Exam — Center for AI Safety and Scale AI collaborators; retrieved 2026-07-17.
