Nexus RaidOps evaluation methodology
How Scores Work
The simplest way to think about it is that NRO asks six different questions about the same player. It does not take one Warcraft Logs parse and rename it six times.
Each category produces a percentage-style score from 1 to 100 using the evidence that belongs to that category. Score and evidence confidence remain separate, so limited information does not silently become poor performance.
Evaluation category
Role Performance
Question: How well did the player perform the combat job they were brought to do?
We use Warcraft Logs information such as:
- Damage, healing, mitigation, or support output
- Overall and item-level parses
- Boss and priority-target contribution
- Cooldown timing
- Rotation, resources, and active time
- Performance during important encounter windows
- Role-specific survivability
- Gear, assignments, strategy, and kill-time context
A DPS player's score leans heavily on useful damage and priority targets. A healer's score emphasizes healing when it mattered, efficiency, cooldowns, and death prevention. A tank's score emphasizes survival, mitigation, control, and stability.
A high parse helps, but it does not automatically produce a high Role Performance score. Padding, dying early, ignoring priority targets, or misusing cooldowns can lower the contextual score.
Progression pulls do not always have a usable Warcraft Logs ranking. When a ranking is missing or comes back as an unranked zero, the boss report uses the player's output percentile among players doing the same role on that boss, along with active time and boss-target contribution. An unavailable ranking is never treated as a real 0 score.
Evaluation category
Utility and Team Contribution
Question: How well did the player use the extra tools available to them to help the raid?
This covers things outside their normal rotation and outside specifically assigned jobs, such as:
- Unassigned interrupts, dispels, purges, and crowd control
- Emergency externals, rescues, grips, or off-healing
- Good battle-resurrection decisions
- Covering another player's missed responsibility
- Helping recover a pull
- Useful strategy information or preparation support
Those opportunities are weighted. An external that prevents a death matters more than an optional crowd-control cast that barely changes anything.
We compare players against the tools their actual class, specialization, and talent build had available. A class with fewer utility abilities is not automatically punished, and a class with many abilities does not receive free credit just for having them.
Attendance by itself is not a Utility failure. When verified utility actions exist, they drive 70% of the score and boss contribution context supplies 30%. When the current log has no recognized utility actions, NRO uses only that boss's activity, survival, and role-relative contribution so the report can still show a fair score instead of 0 or N/E.
Internal weighting
Evaluation category
Consistency
Question: If we put this player in the same situation again, can we expect roughly the same level of performance?
We use every recorded pull and kill from this boss across the Team Season. Normal, Heroic, and Mythic are calculated separately, so the comparison never mixes difficulties. Within that lane, we account for:
- Same boss and difficulty
- Similar progression stage
- Similar phase reached
- Same role, specialization, and general build
- Similar assignments and strategy
- Similar opportunity to perform
We then look at:
- How often they met expectations
- How much their results moved between pulls
- Their normal performance floor, not just their average
- Repeated mechanical or assignment problems
- Unexplained low-performance pulls
- Whether the same stability continued across raid nights
This means two players can average the same result but receive different Consistency scores. Someone scoring 90, 88, 91, and 89 is more consistent than someone scoring 50, 99, 55, and 98, even if their averages are similar.
Consistency is not whether someone is good. It is whether their current level, good or bad, is dependable.
When the season has two or more same-boss, same-difficulty pulls, Consistency is 65% typical pull deviation and 35% high-end deviation. If the season only has one valid pull, NRO uses activity at 45%, survival at 35%, and mechanic stability at 20%. That one-pull value is pulled toward 70 and shown as limited evidence.
Internal weighting
Evaluation category
Assignment Execution
Question: Did the player complete the specific jobs the raid plan gave them?
This includes documented or leadership-confirmed assignments such as:
- Interrupts
- Soaks and immunities
- Dispels and crowd control
- Healing or defensive cooldowns
- Tank swaps and positioning
- Priority targets
- Baits and special positioning
- Add control
- Combat-resurrection priority
Assignments are weighted by importance:
We also account for timing, difficulty, whether the player was primary or backup, whether the assignment was actually possible, and whether another failure prevented them from completing it.
A documented or confirmed assignment always takes priority. NRO should not invent a missed cooldown, interrupt, or soak that the player was never clearly given.
Boss reports read the active cooldown plan directly. With a plan, execution is 75% and on-time use is 25%. Without one, the report uses only this boss's valid mechanic responsibilities at 60%, survival at 25%, and avoiding an early progression death at 15%. This fallback measures general boss-responsibility execution, not a made-up cooldown assignment.
Evaluation category
Adaptability and Improvement
Question: When something changed or the player received useful feedback, did they adjust and retain the correction?
We compare performance before and after things like:
- Coaching
- A strategy change
- A new position or assignment
- A cooldown-timing change
- A repeated mechanic problem
- A role, specialization, or build change
We consider:
- How quickly the change was understood and applied
- Whether performance actually improved
- How many reminders were required
- Whether the correction lasted
- Whether the player could apply the lesson to similar situations
Internal weighting
A player does not need to start out poorly to score well. A strong player can earn a high score by quickly adopting new strategies while maintaining strong performance.
The season trend uses every valid observation for the same boss and difficulty, ordered across the full Team Season. It compares the earlier and later halves of that history. Improvement accounts for 70% of the season trend, while the later success rate accounts for 30%. Normal, Heroic, and Mythic never share a comparison pool.
The same-fight signal is always kept, even when season history exists. It checks whether early mistakes stopped or repeated later, then adjusts for repeated failures, near-deaths, and killing blows. When season history exists it supplies 25% of Adaptability; without season history it remains the limited fallback. Evidence from a different boss or difficulty is never used.
Evaluation category
Mechanic Parse
Question: Of the mechanics the player had a fair opportunity to handle, how many did they handle correctly?
We then adjust for:
- Whether the mechanic was actually avoidable
- Whether the player was personally responsible
- Severity of the mistake
- Whether it was one mistake or a recurring problem
- Progression stage
- Strategy instructions
- Evidence confidence
- Whether the mistake caused deaths, downtime, or a wipe
Important protections include:
- Unavoidable damage does not count as failure.
- Damage caused by another player should not punish the victim.
- Intentional strategy damage is excluded or adjusted.
- Multiple damage ticks from one mistake count as one incident.
- A death is not automatically a mechanical failure.
- A critical error matters more than harmless momentary contact.
The Mechanic Parse is not simply avoidable damage taken, and it is not a Warcraft Logs ranking against other players. It is an opportunity-based measurement of encounter execution.
Reading the result
What the final scores mean
All six categories use the same percentage-style 1–100 scale, while each category uses its own evidence and calculation.
85–100
Exceptional
70–84.9
Strong
55–69.9
Solid
40–54.9
Developing
Below 40
Needs Attention
N/E
Not enough fair evidence