A dashboard computed from your own beliefs is not information
The test for any instrument on a capital programme is whether its output contains something the inputs did not. A rule applied, a calculation done, a source consulted. A rating a team gives itself contains none of these, and colouring it does not add one. One public dashboard has carried both kinds side by side since 2009, and an auditor has re-rated the self-assessed column and published the result.
60 of 95
self-assessed risk ratings on the federal IT Dashboard where GAO's own assessment of the same investment showed more risk, April 2015
On 4 September 2026 this publication scrapped a tool it had built the day before. The tool asked 11 questions, one for each activity LEED's commissioning prerequisite requires, and showed 11 ticks. It looked like an instrument. It was a mirror. The reader knew whether the owner's project requirements existed before they arrived, told the page, and the page told them back with a tick beside it. The output contained nothing the inputs did not.
That sentence is the test this article proposes for every dashboard, status report and readiness check on a capital programme, and it is worth stating carefully because most of them fail it without anybody noticing. An instrument earns the name when its output holds something the person entering the inputs did not already have. There are three ways that can happen, and only three that we can find.
The first is a rule applied. The rebuilt commissioning tool asks who employs the commissioning authority and whether they sit on the project team, and returns where LEED's published tiers place that arrangement in the edition chosen, or that the text does not decide it. The reader knew who they had hired. They did not necessarily know that LEED v4 admits a firm employee who is not on this project's team at the prerequisite tier and excludes one who is, that its small-project relief switches off at 1,860 square metres, or that LEED v5 removed that relief and stopped naming who may serve at all. The rule did the work, and the rule is somebody else's, which is what makes the output checkable.
The second is a calculation done. The same tool takes the reader's floor area and construction cost and applies the LBNL medians to both, which produces two figures for what commissioning should cost on this project. They disagree unless the project costs exactly what the median project cost per unit area, and the size of the disagreement is the information: it says how far this building is from the sample before either number is quoted at anybody. Neither figure was in the reader's head, and the arithmetic is on the page.
The third is a source consulted. A question bank drawn from a gate review tells the reader what the reviewer will ask. That is real information, because the reviewer's questions are not the team's questions, and the gap between them is where gates are failed. What the bank cannot do is answer. The answers are still the reader's, and a page that adds them up has added up beliefs.
A rating a team gives itself has none of the three. Nobody applied a rule, nothing was calculated, and the only source consulted was the party being measured. This is a property of the construction rather than a view of the people: every term in a total of self-ratings is a judgment already held by the person reading the total, so nothing can enter the output that was not in the inputs, whatever the weights. Making it red, amber or green does not change that. It changes how it looks.
A dashboard that has carried both kinds for 17 years
The United States federal government has run a public dashboard of its major information technology investments since June 2009. By 2016 it listed over 770 investments at 26 agencies, covering 42bn dollars of a planned 82bn in annual spending. It is useful here because it displays two kinds of rating side by side, and because the Government Accountability Office has audited it repeatedly and published what it found.
The first kind passes the test. The rule is published, the thresholds are numbers, and a reader with the underlying data can recompute the colour. GAO's earlier audits of the dashboard, in 2010 and 2011, found problems with the accuracy and reliability of the data in that column too. But those are data problems, and a data problem can be found by comparing the colour with the rule. The second kind has no rule to compare against. The only way to audit a judgment is with another judgment, which is what GAO eventually did.
Two things have to travel with that number. GAO built its own assessment from the agencies' risk registers, cost and schedule data and review board briefings, and says so, which means it was reading the same documents the CIOs had and scoring them by a stated method rather than measuring anything independently. And GAO states that CIOs could have had more information than it examined, and that risk ratings are by nature judgmental. The finding is not that 60 chief information officers were wrong. It is that when a second party applied a published method to the same evidence, the self-assessed colour came out lighter than the method did in almost two thirds of cases, and there was no way to tell which cases from the dashboard itself.
The report GAO wrote four years earlier explains part of why. It examined six agencies' ratings from the dashboard's launch to March 2012 and found that CIOs had rated a majority of investments low or moderately low risk throughout, that two agencies had rated nothing high or moderately high risk in that whole period, and that about 47% of investments had received the same rating in every single period, which is the stillness another article here treats as a finding in itself. One agency's process was candid about a reason.
Officials at another agency told GAO in 2016 that programme managers may score risks higher to flag an issue for management attention. That is the same phenomenon in the other direction: the number is being used to send a signal, and a number that carries a signal is not carrying a measurement. Both are entirely reasonable things for a manager to do. Neither is what the dashboard says the cell contains.
Which reading of the divergence
A requirement that is widely unmet can mean several things, and this publication's rule is to say which one it takes. Here the reading is narrow. It is not that self-assessment is dishonest, and GAO does not say so either; it lists process causes, stale updates, slow cycles, factors that ignored active risks, and attributes the rest to judgment. The reading is that a self-rating is an input, and the dashboard presented it as an output. It sat in the same table as the computed cost and schedule colours, in the same three colours, with nothing to tell a reader that one column had been calculated from a rule and the other had been decided by the person the rule would have been applied to.
The consequence is visible in how the ratings were then read. GAO's record of its 2012 recommendation notes that OMB began analysing the ratings in its budget documents, and that in the fiscal year 2017 document OMB described the share of investments self-rated low or moderately low risk rising from 69% in 2012 to 77% in January 2016 as showing continued improvement in the general health of IT investments. That is a dashboard of beliefs being read as a dashboard of the world. It may also have been true. The point is that the number cannot tell you, and a number that cannot tell you whether the world improved or the rater relaxed is not information about the world.
What follows for a capital programme
Every capital programme of any size has this table. It is the monthly status report with a red, amber or green against each workstream, the readiness tracker with a percentage complete against each handover condition, the risk register with a probability and impact assigned by the risk owner. In each, the delivery team is the source of the rating on the delivery team, and the column is coloured the same as the columns computed from the schedule.
Applied to a readiness tracker, most cells turn out to be sourced at best. The gate criteria are published, so the questions are real. The answers are the receiving organisation's own account of itself, and a percentage across them is a count of that account. It says how much of the checklist the organisation says it can show. It does not say how ready the organisation is, and a page that labels it readiness has labelled a belief.
This publication's own instruments
The test applies here. When this article was published, 11 of the site's 14 tools computed, applied a published structure or consulted a source, and each said which. Three were assessments, and they are the ones to be honest about.
The operational readiness check draws every question from a source, a clause of ISO 55001, a gate criterion or a section of a national manual this site indexes, and it applies one rule: a no to any of five questions is reported as a blocker whatever the rest of the answers say, because an operator without permits in its name is not ready at any percentage. That part is computed. The percentage is not. It is the reader's own answers, counted, with equal weights because no source weights them. Until this article was published the tool's page did not say so. It now does, and the number is labelled as what it is: a count of what the reader says can be shown, not a measure of readiness.
The reporting history check refuses to produce a colour at all, for the reason this article gives: it counts how often the ratings on a programme have ever changed, and reports the count without grading it, because a grade on that count would be one more belief with a colour. The commissioning tool, after its first day, computes.
None of that makes the tools better than a status report. It makes them labelled. A board that wants the same from its own dashboard can ask one question of every cell: who produced this, and what would it take for it to be wrong. Where the answer is the team being measured and nothing, the cell is a belief, and it should look like one.
Read the sources
- GAO, IT Dashboard: Agencies Need to Fully Consider Risks When Rating Their Major Investments, GAO-16-494, June 2016GAO-16-494, June 2016; read 2026-09-05.Free in full. Compares 95 self-assessed CIO risk ratings with GAO’s own assessment of the same investments; appendix I gives the method and appendix III lists every pair.
- GAO, Information Technology Dashboard: Opportunities Exist to Improve Transparency and Oversight of Investment Risk at Select Agencies, GAO-13-98, October 2012GAO-13-98, October 2012; read 2026-09-05.Free in full. Six agencies, June 2009 to March 2012. The recommendation status on the same page records OMB’s later reading of the ratings.
- USGBC, LEED v4 BD+C, Fundamental Commissioning and Verification prerequisiteLEED v4 BD+C; read 2026-09-05.The v4 prerequisite text in full, including who may serve as the commissioning authority. Superseded as the current edition by LEED v5 (below); kept because v4 states the employer tiers in its own words and v5 does not.
- Lawrence Berkeley National Laboratory, Building Commissioning: A Golden Opportunity, 20092009; read 2026-09-02.Free. The 2009 meta-analysis: 643 buildings, costs, savings and paybacks.
Related reading
An extension of time is tested against records made during the delay
A claim for more time turns on three separate questions. Was the contract's notice procedure satisfied? Is the event one for which the contract allows an extension, or whose risk it otherwise places on the employer? Did it actually delay completion? Money is a separate claim, with its own proof of causation and amount. Each question is later tested against what was written down while the event was happening, by whichever party has to prove the point. Records do not create entitlement, a record of an event does not prove that it delayed completion, and the prevention principle does not rescue a claim under every contract or every governing law.
ReadThe score that is not a number
A technical score on a qualitative criterion records what evaluators judged. The weighting decides how much of the award that judgement controls; it does not turn the judgement into a measurement. What makes the score answerable is a record that a reader outside the room can test against the solicitation and the proposal, and the instruments examined here require different parts of that record.
ReadEvery award rule makes you publish the weighting. None makes you defend it.
The WTO agreement, the UNCITRAL Model Law and the EU directive all require the relative importance of the evaluation criteria to appear in the tender documents. None says what the weighting should be, and the Model Law's own commentary calls it discretionary. All three specify the arithmetic completely in exactly one situation, an electronic auction, where a machine does the ranking. Meanwhile a 70/30 split under the World Bank's own price formula prices one technical point at 3.4% of the contract and puts a 30 point quality lead beyond the reach of any price at all.
Read