A mayor announces 10,000 new affordable homes. An agency issues a press release calling the effort “on track.” Six months later, residents still need to know the operational facts: How many homes have actually started construction? How much funding has been committed? Which deadlines have moved? How to monitor agency performance starts by treating announcements as the beginning of oversight, not evidence that a result has been delivered.
For city residents, journalists, and policy professionals, the problem is rarely a shortage of information. It is fragmentation. Performance evidence sits across budget books, procurement notices, City Council testimony, executive orders, agency dashboards, audit reports, and public records requests. A useful monitoring system brings those records into one view and applies the same standards over time.
Start with the agency’s actual mandate
An agency should be evaluated against responsibilities it can reasonably control. That sounds obvious, but public debate often skips this step. A transit agency may be blamed for a service disruption caused by a separate construction authority. A housing department may receive credit for a project that still depends on financing, land use approval, or a developer’s execution.
Begin by defining three things: the agency’s legal or administrative mandate, the program or promise being tracked, and the decision-maker responsible for delivery. City charters, enabling laws, executive orders, budget documents, and agency strategic plans can establish this baseline.
This distinction protects against two common errors. First, it prevents an agency from taking full credit for outcomes driven largely by outside forces. Second, it prevents observers from labeling a program “broken” when the responsible agency has completed its assigned step and another institution is holding up the work.
A performance tracker should identify ownership plainly. If responsibility is shared, say so. If a mayoral directive requires action from multiple agencies, track each required action separately rather than assigning one broad status to the entire initiative.
Set a baseline before judging results
Performance cannot be measured against a press release alone. Every claim needs a starting point.
If an agency promises to reduce permit processing times, record the prior median processing time, the number of pending applications, the stated target, and the deadline. If it promises to add mental health services, establish the existing number of providers, locations, appointment slots, and patient wait times. If it commits to cleaner streets, clarify whether the standard is sanitation inspections, resident complaints, litter index scores, or something else.
A baseline turns a political claim into a testable proposition. It also exposes vague commitments. “Improve service” is not a performance standard. “Reduce average wait time from 45 minutes to 30 minutes by the end of the fiscal year” is.
Not every agency will publish the necessary baseline itself. In those cases, use the most authoritative available record and label any limitation. A Comptroller audit may provide a stronger starting point than an agency spokesperson’s estimate. Administrative data can be useful, but it may reflect how the agency counts activity rather than what residents experience.
Choose indicators that show delivery, not activity
A public agency can be very busy and still fail to produce the result it promised. That is why monitoring should separate inputs, outputs, and outcomes.
Inputs are resources: money appropriated, staff hired, contracts awarded, or equipment purchased. Outputs are direct products of agency work: inspections completed, applications processed, shelters opened, trees planted, or cases closed. Outcomes are the lived effect: fewer unsafe buildings, shorter waits, more stable housing, or lower injury rates.
All three matter. Inputs show whether the agency has the capacity to act. Outputs show whether work is occurring. Outcomes show whether the work is making a difference. But they should never be treated as interchangeable.
A sanitation department might purchase new trucks and complete more collection routes. Those are meaningful inputs and outputs. If missed-pickup complaints do not decline in the neighborhoods the program was meant to serve, the outcome remains unresolved. The program may be underpowered, poorly targeted, or simply too new to assess. The appropriate status is not automatically “failed,” but it is not “kept” either.
For most agency initiatives, a compact scorecard should track four to six indicators at most. More metrics can create the appearance of rigor while making the central question harder to answer. Select measures that a resident can understand and a reporter can verify.
Build an evidence file, not a clip file
The strongest monitoring work relies on primary records whenever possible. News coverage can identify a development worth tracking, but it should not be the sole proof that an agency met a target.
Useful evidence commonly comes from four places:
- budget documents and financial management reports, which show appropriations, actual spending, vacancies, and program changes;
- agency reports, dashboards, and datasets, which show operating activity and stated performance measures;
- oversight records, including Council hearings, Inspector General reports, Comptroller audits, and court filings; and
- procurement and implementation records, including contracts, solicitations, project schedules, and change orders.
Each item should have a date, source, and specific relevance to the claim under review. This is less glamorous than collecting quotes, but it makes the assessment reproducible. A reader should be able to see why a promise was marked kept, broken, stalled, or still in progress.
Source discipline also means recording contradictions. Agencies sometimes revise a number, change a methodology, or stop publishing a measure that was previously prominent. Do not quietly replace the old figure with the new one. Note the revision and explain what changed. A disappearing metric can be an accountability item in its own right.
Use statuses that reflect reality
Binary pass-fail judgments are useful only for simple commitments. Government delivery is often staged, delayed, partially completed, or dependent on another approval. A more honest system uses clear status definitions.
“Kept” should mean the agency completed a specific commitment and the available evidence supports that finding. “Broken” should mean the deadline passed or the commitment was abandoned without completion. “Stalled” should mean material progress has stopped, deadlines have repeatedly moved, or a necessary action remains unfulfilled. “In progress” should be reserved for commitments that remain within a plausible timeline and show documented movement.
“Partially kept” is appropriate when a measurable component has been delivered but the full commitment has not. For example, an agency that opened three of five promised service centers should receive credit for the three, not the five.
Status labels need written rules. Without them, the monitoring system becomes vulnerable to mood, ideology, or the latest headline. The point is not to make government look better or worse. It is to make the record legible.
Monitor the budget as closely as the promise
A program’s budget is often the clearest early warning signal. A mayor can announce a major initiative, but if the funding is absent, delayed, reallocated, or left unspent, the promise may not be operational.
Track the difference between proposed funding, adopted funding, modified funding, and actual spending. These are not the same thing. A budget line may survive adoption but be reduced during the year. An agency can receive funds and struggle to hire staff. A capital project can be funded but delayed by procurement, permitting, or design changes.
Vacancy rates deserve similar attention. Agencies frequently cite staffing shortages, and sometimes that explanation is valid. But the public should be able to distinguish between a citywide hiring constraint and an agency that has failed to fill authorized positions or move a hiring process forward.
Budget scrutiny requires patience. Some programs are seasonal, and capital spending is irregular by nature. A low spending rate in the first quarter is not automatically a red flag. The relevant question is whether spending and staffing match the project schedule the agency itself established.
Check for distribution, not just citywide totals
Citywide averages can conceal unequal service. An agency may report that average response times improved while particular neighborhoods, languages, disability communities, or housing types saw no improvement at all.
Where data allows, break out performance by borough, neighborhood, demographic group, service type, and time period. This is especially necessary in housing, education, health, public safety, sanitation, and transportation, where access and burden are rarely distributed evenly.
Disaggregation has limits. Small samples can produce unstable rates, and privacy rules may restrict what can be published. Those constraints should be stated, not used as a reason to avoid the question. If the data is insufficient to assess equitable delivery, that is a finding worth reporting.
Establish a regular review cycle
Agency performance changes slowly enough that daily judgment is usually noise, but quickly enough that annual reviews can arrive too late. The right cadence depends on the program.
Operational measures such as 311 response times, shelter utilization, permit backlogs, or service interruptions may warrant monthly checks. Budget execution and hiring are often best reviewed quarterly. Long-term capital projects may need milestone-based reviews tied to design completion, procurement, construction start, and opening dates.
Every update should answer the same questions: What changed? What evidence supports the change? Did the target, timeline, funding, or ownership change? What remains unverified? Consistency is what turns scattered updates into public oversight.
A live accountability dashboard can make this work easier to follow, but the format is not the accountability. The method is. A clean scorecard still needs definitions, dated evidence, and a willingness to revise a finding when the record changes.
The public does not need to become a budget analyst to hold government accountable. It needs a clear claim, a credible baseline, evidence of delivery, and a record of what happened when deadlines arrived. That is enough to replace political theater with a question every agency should be prepared to answer: what did you say you would do, and what can you prove?
