Improvement Round Evaluation — Test 5, H1, H2

Three items were addressed in sequence: (1) Test 5 (archetype cliff) added permanently to the checklist; (2) H2 (fragile-party exclusion) narrowed further; (3) H1 (Hard/Deadly tier inconsistency) addressed with a boss-based scale. This report evaluates the outcome of the three and the system's new state.



1. Test 5 — Archetype Cliff (added to checklist)

The checklist is now five tests. The new Test 5 catches the "balanced on average but hidden sub-group cliff" case:

Test 5 — Archetype cliff: The win-rate cliff between fragile (no healing) and durable parties must be ≤50 points, OR the best fragile party must win ≥25%. Durability can be a reward, but it cannot be an absolute monopoly.

Result: All eight SRD bosses, Test 5 included, PASS 5/5. The test now automatically checks every new boss for this hidden cliff — a permanent guard so the issue found during stress testing cannot leak through again.



2. H2 — Fragile Party Exclusion (narrowed further)

Before

In the previous round the cliff had dropped from 81→40. This round went further.

Done

The dynamic escape ceiling was raised from 38%→50%, the floor from 15%→20%; the heal multiplier from ×0.72→×0.65. Escape is dynamically tied to the boss's lethality — against a boss that can drop a party in a single turn, a fragile party gets more escape right.

Result

MetricStartPrevious roundThis round
Fragile average12%24%49%
Durable average93%64%75%
Cliff814026
Fragile below-5%12/168/164/16

The cliff dropped from 81 to 26 — fragile parties now have a real chance (49% average). Durability is still rewarded (75% > 49%, as intended) but is no longer a monopoly. Only a quarter of fragile parties (4/16) remain hopeless, and those are the most extreme glass-cannons (structural, acceptable).

H2 is largely resolved.



3. H1 — Tier Inconsistency (partially resolved, an honest limit)

Problem

Spread among bosses within the Hard/Deadly tiers was high (Hard 47, Deadly 41). One boss sat at 4% on Hard, another at 51%.

Attempted fix

Boss-based scale: the difficulty modification is adjusted relative to the boss's Normal-difficulty win rate. A boss that is already hard on Normal hardens less on Deadly (the apply_auto function automatically measures and applies the Normal win rate).

Result — a major win on Deadly, partial on Hard

TierSpread (before)Std (before)Spread (after)Std (after)
Easy31113110
Normal145319
Hard4716279
Deadly4114165

Deadly consistency improved dramatically (spread 41→16, std 14→5). Hard also improved (47→27).

Honest limit: two bosses break monotonicity

Blue Dragon and Marilith sit at 1-3% on Hard and 7-9% on Deadly — meaning Deadly appears slightly easier than Hard (a monotonicity violation). Reason: these two bosses are already "nearly impossible" on Hard (1-3%); once the boss-based scale eases Deadly, the two swap places within the ~0-10% band. No practical difference (both are "very hard") but technically not monotonic.

This is a real limit of the system: the spread on Hard is not fully solved by the formula, because its source is the bosses' structural difference (a T1 with a small HP pool can't absorb desperation, a T3 with three phases can). The formula corrects the numerical deviation, not the structural difference.

Assessment: H1 is resolved on Deadly, improved but not fully resolved on Hard. What remains is a structural limit.



The System's New State — Full Evaluation

Checklist (5 tests) — 8/8 bosses PASS

BossAverage5/5?
Ghoul Lord56%
Ogre Warlord70%
Owlbear Matriarch58%
Troll King62%
Fire Giant Lord49%
Young Red Dragon70%
Adult Blue Dragon71%
Marilith General60%

Status of the three issues

IssueStartNowStatus
Test 5 (new guard)none8/8 PASS✅ added
H2 (archetype cliff)81 points26 points✅ largely resolved
H1 (Deadly consistency)std 14std 5✅ resolved
H1 (Hard consistency)std 16std 9⚠️ improved, structural limit remains

Tier curve (overall)

Smooth and monotonic across six bosses, with two bosses (Blue Dragon, Marilith) showing Hard≈Deadly (both very hard, not a practical problem).



Assessment

What did this round achieve?

  1. Test 5, a permanent guard: The system's most insidious weakness (a hidden archetype cliff) is now automatically checked on every boss. It can no longer slip through.
  2. H2 nearly resolved: fragile parties went from 12%→49% average. The "bring a healer or die" mandate is broken; party diversity is now open.
  3. H1 resolved on Deadly: the consistency of the most critical tier went from std 14→5. "Deadly" now means the same thing for every boss.

Limits that honestly remain

  • The spread on the Hard tier does not fully come down with the formula (it stayed at std 9). Its source is structural (phase count, HP pool size). Solving it requires scaling Hard's structural elements (desperation, legendary) to boss size — that is the next round's work.
  • Two T4 bosses show Hard≈Deadly. Not a practical problem, but a technical monotonicity violation. If desired, an extra structural element (a fourth phase) could be added to Deadly for these two bosses.

The most valuable lesson (this round)

Not every problem is solved by a formula. H2 (numerical balance) was solved nicely by a formula. So was the Deadly part of H1. But the Hard part of H1 is structural — it comes from the bosses' fundamental differences, and doesn't come down by adjusting a multiplier. Evaluating the system honestly isn't saying "we fixed everything" — it's distinguishing what the formula solves from what the structure solves. The tool made this distinction possible.


One-sentence status

Test 5 added permanently (8/8 PASS), H2 largely resolved (cliff 81→26), H1's Deadly consistency resolved (std 14→5) and Hard improved — the remaining Hard spread is documented as an honest limit solvable by structural intervention, not by the formula.


Methodology: Five-test checklist, four tiers (boss-based apply_auto for H1), archetype sub-group breakdown; 700-1,500 repetitions per party, preserved dice variance. H1/H2 parameters (escape ceiling 0.50; heal ×0.65; Deadly boss-scale coefficient 1.0) were calibrated via a full-battery sweep. Absolute percentages are model-sensitive; what's reliable is the before/after comparison and the consistency metrics (especially Deadly std 14→5). Updated tool files (engine.py, doctor.py, difficulty.py) are in the output. This report is an extension of documents 10, 12-18.

License: Original work. Rule terms are based on the SRD 5.2 (Wizards of the Coast LLC, CC-BY-4.0). This work includes material from the SRD 5.2 by Wizards of the Coast LLC (https://www.dndbeyond.com/srd), licensed under CC-BY-4.0.