Skip to main content
Workflow & Fixes

Triage Bugs or Fix Them? A Practical Workflow

Every bug that lands in your queue is a decision: dig in, or bounce it back? We think triage belongs in the loop. Just not the way most teams run it.

A bug shows up in your queue. "App crashes when uploading large files." No repro steps. No environment. No idea whether it's a one-off or the thing that kills your Friday.

Now what?

That pause — investigate or send it back — is triage. Most teams skip it. That's usually a mistake, not because triage is holy, but because without it your senior engineers spend their mornings doing data entry for a bug tracker.

What follows is an argument for a real triage step. Where it earns its keep, how to keep it from eating your week, and the cases where it genuinely doesn't matter.

What triage actually buys you

Bug trackers fill up with junk. Duplicates. Feature requests wearing a bug costume. Reports that say "it's broken" and stop there.

If every new report goes straight to a developer, you're paying senior-engineer money for clerical work.

Atlassian's bug life cycle sketches Open → In Progress → In Review → Done → Closed, with triage as an optional detour. That optional detour is the whole ballgame. Someone looks at each new bug, checks whether it's real and complete, sets priority and severity, then accepts it or bounces it back with a reason.

Concrete case. A team I know was pulling roughly 60 reports a week. Before triage, developers burned about 22 hours investigating stuff that turned out to be duplicates or not bugs at all. They added a rotating triager. That person spent 6 hours a week filtering. Net gain: about 16 developer hours back, every week. Call it a 73% cut in wasted effort if you like round-ish numbers. The point is the 6 bought back the 16.

Time isn't the only thing. Priority and severity are different animals, and teams conflate them constantly. Priority is urgency. Severity is impact. A crash hitting every user is high severity. If it only fires inside a feature nobody opens, priority can sit low for months. Triage forces that call early — before someone sinks a day into the wrong bug.

Running triage that doesn't make you miserable

Good triage is a habit, not a heroic weekly push.

Define "ready" in writing. A bug is triage-ready when it has repro steps, expected vs. actual behavior, environment details, and a severity guess from the reporter. Missing any of those? Bounce it back with a template. The NVD works the same way — it doesn't test vulnerabilities itself, it leans on vendors and researchers to supply the data. Your triager leans on reporters.

Name your exception states. Open/In Progress/Done isn't enough. You need Rejected, Duplicate, Deferred, Reopened, Not a Bug, Non-Reproducible. "Non-Reproducible" doesn't mean fixed. It means parked until someone produces evidence. Triage is where those labels get applied — not six weeks later in a retro.

Pick a scale. Any scale. Then never change it. Priority from Lowest to Blocker. Severity from Minor to Fatal. For security work, CVSS maps scores to ratings: None (0.0), Low (0.1–3.9), Medium (4.0–6.9), High (7.0–8.9), Critical (9.0–10.0). The labels matter less than using them identically every time. Teams that let "High" drift end up with a backlog where everything is High, which is the same as nothing being High.

Get more than one brain in the room. A developer, a tester, and a product owner catch different things. The developer sees technical dependencies. The product owner knows what a customer will scream about. Triage as a solo chore is triage that gets skipped the first busy week.

One thing I'd add that most triage guides skip: timebox the triager's investigation. Fifteen minutes per bug, hard stop. If it isn't clear in fifteen minutes, it isn't a triage problem anymore — it's an investigation task, and it needs its own ticket. Without that rule, your triager becomes a part-time detective and the queue backs up anyway.

When triage fails, and when it doesn't matter

Skipping triage can be expensive. The Therac-25 accidents between 1985 and 1987 killed and injured patients, and the investigation found most computer-related accidents traced back to requirements errors — omissions, mishandled environmental conditions, bad system states — not just coding bugs. A triage process that questions assumptions might have caught some of that earlier.

But I'll be honest: triage alone wouldn't have stopped Therac-25. That was a systems failure. Blaming it on bug tracking misses the point.

Same with Ariane 5 Flight 501. A 64-bit float got converted to a 16-bit signed integer, blew past 32,767, and the rocket came apart. The inquiry board pointed at specification and design errors. Triage for requirements, not just bugs, might have helped. Might. No silver bullet.

And here's the part triage evangelists leave out: not every bug gets more expensive when you fix it later. A study of 171 projects from 2006 to 2014 found no evidence that issues resolved later cost substantially more effort. The delayed-issue effect shows up intermittently, in some projects only. So triage isn't about fixing sooner. It's about fixing the right things. If you know which bugs punish you for waiting, you prioritize those and let the rest sit.

Security is the exception. There, triage isn't optional. CISA's Known Exploited Vulnerabilities catalog lists remediation due dates and flags whether a vulnerability is used in ransomware. BOD 22-01 forced federal agencies to fix KEV vulnerabilities; it's since been superseded by BOD 26-04. If you're not triaging security bugs with that kind of rigor, you're gambling with someone else's data. The 2024 CWE Top 25 puts CWE-79 (Cross-site Scripting), CWE-787 (Out-of-bounds Write), and CWE-89 (SQL Injection) at the top. Triage is how you map incoming security bugs to those categories and sort by exploitability instead of by whoever shouted loudest.

Lightweight or heavyweight?

Not all triage looks the same. A daily 15-minute standup is one shape. A formal weekly meeting with a dedicated triager is another. Which one fits depends on your team, your bug volume, and how much a bad miss would hurt.

Criteria Lightweight Heavyweight
Frequency Daily, 15 min Weekly, 1 hour
Participants 2-3 team members 5+ including product, dev, QA
Bug Volume Low to medium (<20/week) High (>50/week)
Risk Profile Non-critical apps Safety-critical or security-sensitive
Tools Basic tracker Advanced tracker with custom fields

Most teams can survive on the lightweight version. If you're shipping medical devices or spacecraft software, you need the heavier one. NASA JPL's Gerard Holzmann proposed "The Power of 10: Rules for Developing Safety-Critical Code." Triage in that world includes checking compliance against those rules.

There's also the option of static analysis ahead of triage. It finds defects without running the program. CodeSonar, adapted with NASA JPL SBIR funding, is used by hundreds of organizations, and the FDA has nudged infusion-pump makers toward tools like it. Catch bugs automatically and triage becomes about prioritizing what's left. But static analysis won't tell you business impact. A human still makes the priority call.

Keeping triage from becoming a tax

The loudest complaint about triage is overhead. A good process actually saves time by cutting context switching. A few ways to keep it lean:

  • Automate the boring parts. Route new bugs into a triage queue, set a default priority, ping the triager. Bugzilla supports custom fields and workflow management for exactly that.
  • Rotate the triager weekly. Don't park one person on it forever. Rotation spreads the load and teaches everyone what's actually coming in the door.
  • Timebox the meeting. Fifteen minutes daily covers most teams. Anything needing deep investigation becomes its own task.
  • Enforce the definition of ready. No repro steps, no expected vs. actual, no environment? Back to the reporter. No exceptions, or the definition stops meaning anything.

Triage isn't fixing. It's deciding what to fix, in what order, with enough information to be right more often than you're wrong. The output is a prioritized backlog of well-defined bugs. That's the deliverable.

One more habit worth building: search for duplicates before you accept anything. A thirty-second query can save someone an afternoon. Bugzilla's design principles lean on speed and efficiency for a reason.

Where this leaves you

If your workflow has no triage step, add one. It doesn't need to be elaborate — a daily 15-minute meeting and a clear definition of ready will do. Triage cuts noise, keeps priority and severity consistent, and stops critical bugs from quietly getting lost in the pile.

A lot of software failures come from a small number of variables. Triage is how you find them before your users do.

Sources

  • Atlassian (bug life cycle) - https://www.atlassian.com/software/jira/guides/bug-tracking/bug-life-cycle
  • Atlassian (issue priority vs severity) - https://www.atlassian.com/software/jira/guides/issues/priorities
  • MIT (Therac-25 accidents Part V, Leveson & Turner) - https://web.mit.edu/6.033/2004/wwwdocs/papers/Therac_5.html
  • University of Minnesota (Ariane 5 Flight 501) - https://www-users.cse.umn.edu/~arnold/disasters/ariane.html
  • CWE (Top 25 2024 ranked list, MITRE) - https://cwe.mitre.org/top25/archive/2024/2024_top25_list
  • FIRST (CVSS v4.0 specification) - https://www.first.org/cvss/v4.0/specification-document

Share this article:

Comments (0)

No comments yet. Be the first to comment!