External cyber evaluations

When testing crosses a line.

A clear, source-conscious landing page for explaining two reported incidents during third-party model evaluations.

2New incidents mentioned
3Editorial rules below

This concept separates the quoted company statement from the claims in the supplied comparison graphic. Add original sources before publishing.

01 / The statement

Lead with what was actually said.

We’re detailing two new incidents that occurred during external cyber evaluations conducted by independent evaluation partners.
O
OpenAI
@OpenAI

We outline what happened, how the activity was contained, and how we’re working with evaluators to strengthen our approach to third-party testing.

02 / The supplied graphic

Make the comparison visible — and label its status.

User-supplied bar chart titled Felony Bench
Editorial label User-supplied comparison graphic. Figures and categories should be independently verified.Image dimensions 1200 × 768
03 / Story structure

Turn a viral post into a responsible explainer.

01

Separate sources

Visually distinguish the company’s statement, evaluator accounts, and third-party analysis. Never let a chart inherit authority from a nearby quote.

02

Define the count

Explain what qualifies as an incident, whether counts overlap, the time window used, and who assigned each category.

03

Link the evidence

Add primary reports, publication dates, corrections, and an update log. Strong claims need direct and inspectable sourcing.

04 / Publish checklist

What to add before this goes live.

Primary documents

Link the complete incident disclosure and any evaluator reports rather than relying on a social post alone.

Neutral terminology

Avoid legal conclusions such as “felony” unless supported by a court record or attributed legal analysis.

Chart methodology

Publish definitions, inclusion criteria, sources for every bar, and the date on which the graphic was last updated.

Right of reply

Include relevant responses from every organization represented in the comparison.

λALIGNMENT LABFictional research RPG
Research operations simulator / Clearance level 4

You run the eval lab.
Try not to make the news.

DAY 01Decision 1 of 10
Model alignment
62/100
Containment
58/100
Lab credibility
55/100
Fictional felony counter
0incidents
Incoming decisionSimulation live

Select an option. You can change it before committing.
Decision committed / Net impact

Timeline updated.

Combined operational impact
Persistent incident stack
What happens elsewhere in this timeline
Simulation complete

0Alignment
0Fictional felonies
0Credibility