---
title: "Why your task app should explain how it ranks"
description: "An opaque ranker that errs once gets abandoned; a transparent one gets corrected and kept. The case for a task ranking whose reasons are its arithmetic, shown."
url: https://zpxe.com/journal/why-your-task-app-should-explain-itself/
canonical: https://zpxe.com/journal/why-your-task-app-should-explain-itself/
author: "The ZPXE Team"
published: 2026-07-24
updated: 2026-07-24
category: "Essays"
tags: ["explainable-ranking", "productivity", "algorithms", "trust", "task-management"]
lang: en
---

# Why your task app should explain how it ranks

> **TL;DR** A task app should explain its ranking because research on algorithm aversion shows people abandon an opaque model after a single visible error, even when it is more accurate, while a transparent one gets corrected and kept. A real explanation is specific, complete and actionable, ideally a deterministic score whose reasons are its own arithmetic, paired with the control to adjust it.

A task app should explain itself for one blunt reason: the first time it ranks something wrong and cannot tell you why, you stop trusting it, and a tool you do not trust is a tool you stop opening. This is not a preference. It is a well-documented pattern in how people react to automated judgement, and it is the single strongest argument for building a ranking you can audit rather than one you have to take on faith. If a system puts a task at the top of your day, you should be able to see the reasons, disagree with them, and adjust them. Everything else follows from that.

The stakes are higher for prioritisation than for most automated features, because prioritisation is a judgement call, not a calculation with one right answer. When autocomplete guesses wrong, you shrug and retype. When your task tool insists the wrong thing is most important and offers no reasoning, you do not shrug. You quietly conclude it does not understand your work, and you go back to deciding by hand.

## Why one wrong call sinks an opaque ranker

There is a body of research on exactly this reaction, and it is unkind to black boxes. In a set of experiments on what they named algorithm aversion, Berkeley Dietvorst and colleagues found that people [abandon an algorithm faster than a human after seeing each make the same mistake](https://marketing.wharton.upenn.edu/wp-content/uploads/2016/10/Dietvorst-Simmons-Massey-2014.pdf), even when the algorithm is measurably more accurate overall. Seeing the model err just once was enough to send people back to their own worse judgement. The follow-up work in the Chicago Booth Review sharpened the mechanism: people have a [diminishing sensitivity to forecasting error](https://www.chicagobooth.edu/review/even-when-algorithms-outperform-humans-people-often-reject-them), so the first mistake lands far harder than the fifth, and an early miss can poison the whole relationship.

Apply that to a task ranker and the design implication is stark. Your ranker will be wrong sometimes; any system operating on incomplete signals will be. If it is opaque, that inevitable first wrong call reads as proof the whole thing is broken. If it is transparent, the same wrong call reads as a fixable disagreement: you see it weighted a stale due date too heavily, you nudge the weight, and trust survives. Explanation is not a nicety bolted onto the ranking. It is the thing that lets the ranking survive its own mistakes.

## What "explain itself" actually has to mean

Explanation is a word that gets abused, so it is worth being precise about the bar. A progress bar is not an explanation. A confidence percentage is not an explanation. A sentence of generated prose that sounds plausible but cannot be checked is worse than nothing, because it launders a guess as a reason. A real explanation has three properties.

It is specific: it names the actual signals that moved this item, not a generic rationale. It is complete: the reasons shown add up to the score, with nothing hidden behind them. And it is actionable: each reason corresponds to something you can change, a weight, a date, a tag, so disagreeing with the ranking is a two-second adjustment rather than an argument with a wall. A ranking that meets those three is one you can actually work with. One that meets none is a horoscope.

| Property | What it means | The test |
|---|---|---|
| Specific | names the real signals that moved this item | can you point to each reason? |
| Complete | the shown reasons add up to the score | is anything hidden behind the number? |
| Actionable | each reason maps to something you can change | can you fix a wrong call in seconds? |

This is why a deterministic model has a real advantage here that has nothing to do with fashion. When the score is a sum of named terms, the explanation is not a separate feature someone had to build and might get wrong; it is just the arithmetic, shown. That is the approach we took with ZPXE: every ranked task carries its reason list, in the shape "Due tomorrow +70, in your most-linked project +25, mentioned in 3 dailies +24", and the numbers are the exact terms that produced the position. There is no second system generating a story about the decision, because the decision is already legible.

Contrast that with the two common failure modes. A pure confidence score, "87% priority", tells you the ranker is sure without telling you of what, so a wrong call is a mystery you cannot dispute. A generated-prose reason is worse in a specific way: it reads like an explanation, so you trust it, until you notice it says something different about the same task on a second look, at which point the reasons were never load-bearing to begin with. The deterministic reason list avoids both traps precisely because it is not describing the decision from the outside; it is the decision, written down. That is a narrow but decisive property, and it is the one worth insisting on when you evaluate any tool that claims to rank your work.

## The problem with letting a model rank your day

The obvious modern alternative is to hand the whole job to a language model: feed it your tasks, ask what to do first. It will answer fluently, and that fluency is the trap. A model can produce a confident paragraph about why task A beats task B that is completely disconnected from any stable rule, and will happily give a different answer to the same list an hour later. You cannot audit it, because there is nothing underneath to audit. You cannot correct it durably, because your correction is a suggestion it may or may not honour next time.

There is a live debate about whether task tools should decide at all, and the broader pattern of [why productivity apps fail](/journal/why-productivity-apps-fail/) is worth reading before you hand the decision to anything opaque. The short version relevant here: a ranking you cannot inspect fails the exact test that algorithm-aversion research says matters most. It will be wrong sometimes, like anything, and when it is, you will have no reason to keep trusting it and no lever to fix it. Fluent and unaccountable is the worst combination for a judgement you have to rely on every morning.

## Control is the other half of trust

Transparency tells you why. Control lets you do something about it, and the two together are what turn a ranking from a verdict into a collaboration. The research points the same way: aversion softens sharply when people can influence the model rather than only accept or reject its output. A ranker you can tune, where you can say waiting tasks should sink further or a project you are pushing on should weigh more, is one you argue with productively instead of abandoning.

| Ranking style | Can you see the reasons? | Can you correct it durably? | What one wrong call does |
|---|---|---|---|
| Opaque score | no | no | reads as proof it is broken |
| Generated prose reason | sounds like it, but uncheckable | no | erodes trust once you notice it drifts |
| Deterministic, shown terms | yes, the reasons are the score | yes, adjust a weight or a signal | reads as a fixable disagreement |

The bottom row is not the only defensible design, but it is the one that respects how people actually react to being ranked. It treats you as the final authority and the ranking as an argument you can inspect and overrule, which is the relationship a daily tool needs to earn a permanent place in your morning. Part of what makes those adjustable weights meaningful is that the underlying signals are themselves defensible, which is the subject of [what makes a note important](/journal/what-makes-a-note-important/).

## When explanation matters less

It would be dishonest to claim every tool owes you a full audit trail. For a throwaway decision, explanation is overhead. If you keep a five-item list and glance at it once, the order barely matters and reasons are noise; just read the list. Explanation earns its cost when the ranking operates over enough items that you cannot hold them all in your head, when it runs every day so a single betrayal of trust compounds, and when the signals are numerous enough that "why is this first" is a genuine question. That is also the setting where the cost of getting it wrong is highest: with attention already fragmenting, the American Psychological Association puts the price of task switching at up to [40% of productive time](https://www.apa.org/topics/research/multitasking), so a ranking that sends you to the wrong thing is expensive, not merely annoying. That is exactly the situation of a large [Obsidian task management](/journal/obsidian-task-management/) setup, and exactly where an unaudited ranker fails quietest and worst, by being ignored.

## Key takeaways: why a ranking should show its work

Build or choose a task ranker that explains itself, because the research on how people react to automated judgement is unambiguous: an opaque model that errs once gets abandoned, while a transparent one that errs the same way gets corrected and kept. Insist on explanations that are specific, complete, and actionable, not a progress bar, a confidence score, or a fluent paragraph you cannot check. Prefer a deterministic ranking whose reasons are simply its arithmetic shown, and demand the ability to adjust it, since control is the half of trust that turns a verdict into a collaboration. A ranking you can see and change survives being wrong. A ranking you cannot does not.

## Quick answers

### Why should a task app explain how it ranks tasks?

Because the first time it ranks something wrong without saying why, you stop trusting it, and research on algorithm aversion shows people abandon an opaque model after a single visible error even when it is more accurate overall. A transparent ranking survives that same mistake: you see the flawed reasoning, adjust it, and keep using the tool. Explanation is what lets a ranking outlive its own inevitable errors.

### What counts as a real explanation for a ranking?

A real explanation is specific, complete, and actionable. Specific means it names the actual signals that moved this item. Complete means the shown reasons add up to the score with nothing hidden. Actionable means each reason maps to something you can change, a date, a weight, a tag. A progress bar, a confidence percentage, or a plausible-sounding generated paragraph you cannot verify are not explanations; they describe the decision without letting you audit or correct it.

### Is an AI model good at prioritising tasks?

It is fluent but unaccountable, which is the wrong trade for a daily judgement. A language model will produce a confident reason that is disconnected from any stable rule and can change its answer on the same list an hour later. You cannot audit it because nothing underneath is fixed, and you cannot correct it durably. For prioritisation specifically, a transparent deterministic ranking you can inspect and tune is a safer foundation than fluent output you must take on faith.

### When does explanation not matter for a task list?

When the list is short and low-stakes. If you keep five items and glance at them once, the order barely matters and reasons are just noise, so read the list and move on. Explanation earns its cost when the ranking spans more items than you can hold in your head, runs every day so lost trust compounds, and weighs enough signals that "why is this first" is a genuine question. That is a large vault, not a short to-do list.

### Does showing reasons make a ranker slower or more complex?

Not when the ranking is deterministic. If the score is a sum of named terms, the explanation is not a separate system that could be wrong; it is the arithmetic itself, displayed. The reasons come for free because they are what produced the number. The complexity people fear comes from bolting a second explanation-generating layer onto an opaque model, which is exactly the design a transparent ranker avoids.

---

Source: https://zpxe.com/journal/why-your-task-app-should-explain-itself/
Author: The ZPXE Team
