
Best practices
Measuring impact
Choose what to measure and why.
18 minute read · Reviewed September 2, 2026
What nonprofit impact measurement can establish
Measurement records defined information consistently. Monitoring tracks implementation and change over time. Evaluation uses systematic questions and methods to understand a program, its implementation, outcomes, context, and value for a stated purpose. Outputs are direct products of activities; outcomes are changes in people, organizations, practices, relationships, or conditions. Impact often refers to broader or longer-term change. A program may plausibly contribute to that change without evidence that it caused the change by itself.
Why it matters
- A decision-led plan prevents teams from collecting information simply because it is easy to count, familiar to a funder, or available in a software dashboard.
- Clear definitions make results interpretable across staff, participants, partners, reporting periods, sites, and changes in delivery.
- Combining implementation evidence with outcome evidence helps a team distinguish a weak idea from inconsistent delivery, insufficient reach, changing context, or an inappropriate measure.
- Participant interpretation, disaggregation, access review, data minimization, and explicit limitations reduce the risk of turning incomplete information into a universal claim.
Stage-specific guidance
Measure differently as the organization develops
Begin with context, lived experience, existing evidence, and the people affected. Learn how they define the issue and a meaningful improvement before building a survey or promising an impact target.
- Write the need separately from the proposed service and identify whose perspective, data, or history is represented or missing.
- Draft a small set of possible outcomes with people affected, then test whether each outcome is specific, relevant, observable, and ethically measurable.
- Inventory existing administrative, partner, qualitative, and public information before asking people to provide more data.
Ready to move on when: You can name the intended change, the people affected, the decision the evidence should support, and the perspectives still missing.
Interactive measurement-plan builder
Measurement plan
Choose what to measure and why.
Fictional example
Illustrative example: Willow Street Family Resource Network
A fictional nonprofit is piloting bilingual legal-navigation appointments. Staff want to learn whether the pilot helps residents understand and complete appropriate next steps; they do not control legal-service capacity, agency decisions, court timelines, housing conditions, or case outcomes.
Activity presented as impact
“We served 180 residents and therefore improved housing stability in the neighborhood.”
Bounded outcome and evidence plan
“For pilot participants who consent to follow-up, examine whether understanding of options and completion of an appropriate next step changes within 30 days. Combine appointment records, a brief voluntary follow-up, and participant interviews; report response rate, missingness, access barriers, referral capacity, and alternative explanations.”
The stronger plan separates service volume from participant change, defines who and when, uses complementary evidence, preserves choice, and limits the claim to what the design may support.
A decision-to-use measurement practice
- 01
Name the decision and users
State why the evaluation is being conducted, who needs the findings, when they need them, and the decision they can actually make.
Prompt: Who will use the evidence to improve, continue, pause, fund, report, or expand what—and by when?
- 02
Define the outcome
Name who or what may change, the direction and type of change, an appropriate timeframe, and the relationship to the program pathway.
Prompt: What change is expected, for whom or what, after which experience, and over what period?
- 03
Focus the evaluation question
Ask one open, answerable question aligned with the program stage, intended use, available resources, timing, and relevant equity considerations.
Prompt: What do decision-makers need to understand about implementation, reach, outcomes, variation, or unintended effects?
- 04
Specify the indicator
Define the observable signal, unit, population, calculation, timeframe, exclusions, and interpretation before reviewing results.
Prompt: What exactly will be observed or calculated, and what would it not tell us?
- 05
Choose proportionate evidence
Match sources and qualitative or quantitative methods to the question, credibility needs, feasibility, participant burden, access, and privacy.
Prompt: What existing or new evidence can answer this question with the least reasonable burden and harm?
- 06
Analyze and interpret
Assess quality and missingness, compare evidence with expectations, examine variation, consider alternative explanations, and interpret findings with people who know the context.
Prompt: What patterns appear, how certain are they, who might be missing, and what other conditions could explain them?
- 07
Act and document
Connect findings to a preplanned decision rule, communicate them in usable forms, record the decision and dissent, and define the next learning cycle.
Prompt: What will we continue, adapt, investigate, pause, or stop—and what evidence would change that decision?
Measurement-plan quality checklist
- The plan names the evaluation purpose, intended users, intended use, owner, timing, and decision authority.
- The program pathway distinguishes inputs, activities, outputs, near-term outcomes, intermediate outcomes, and longer-term contribution.
- The outcome states who or what may change, the type or direction of change, and an appropriate timeframe.
- The evaluation question is open, answerable, bounded, useful, feasible, and aligned with the program’s stage and context.
- Every indicator has an operational definition, source, population, timeframe, calculation or coding method, exclusions, and interpretation limits.
- Methods fit the question and credibility needs; qualitative evidence is not treated as anecdotal and quantitative evidence is not treated as automatically objective.
- The plan inventories existing data, minimizes new collection, estimates burden, and defines access, consent, privacy, retention, and deletion practices.
- Collection and interpretation account for language, disability, technology, compensation, power, cultural context, missingness, and safe participation.
- Disaggregation categories are relevant and safe, with minimum reporting thresholds or suppression where groups could be identified or misrepresented.
- The analysis plan identifies expectations, comparison points, missing-data handling, alternative explanations, limitations, and appropriate claim language.
- Participants and relevant partners can interpret, question, or correct findings before consequential conclusions are published or acted on.
- The plan states how findings will be communicated, used, documented, revisited, and protected from overclaiming or selective reporting.
Common measurement failures
- Starting with available metrics instead of a decision.
- Name the intended user, use, and evaluation question first. Keep only measures that can credibly inform that decision.
- Calling outputs impact.
- Report services, sessions, referrals, or materials as outputs. Examine a defined change before making an outcome claim.
- Using a change-over-time result as proof of causation.
- Describe the design, context, comparison, limitations, and alternative explanations. Use contribution language unless attribution is supported.
- Creating a long survey for every participant.
- Inventory existing information, ask only what the decision needs, pilot the method, estimate burden, and provide accessible voluntary alternatives.
- Treating a single average as the whole story.
- Review distributions, missingness, experiences, context, and safe relevant variation without exposing or stereotyping small groups.
- Collecting sensitive information without a lifecycle plan.
- Define necessity, authority, access, security, retention, deletion, consent, risk, and response before collection.
- Publishing only favorable findings.
- Predefine questions and expectations, report null and unintended findings, preserve limitations, and document interpretation and decisions.
Evidence that the measurement practice is useful
Judge the measurement practice by the usefulness, integrity, and responsible use of evidence—not by the number of metrics, dashboards, survey responses, or positive findings.
- 01Decision use: findings are tied to a documented continue, adapt, investigate, pause, stop, budget, or communication decision.
- 02Definition integrity: indicators retain clear definitions, sources, timeframes, calculations, exclusions, ownership, and version history.
- 03Evidence fitness: methods answer the stated question at a level of rigor proportionate to the consequence of the decision and claim.
- 04Participation quality: affected people can shape questions, use accessible voluntary methods, interpret findings, and see how their input affected decisions.
- 05Burden and privacy: requested information is necessary, collection effort is monitored, access is limited, retention is bounded, and avoidable sensitive data is not gathered.
- 06Transparency: reports include response rates, missingness, variation, uncertainty, limitations, alternative explanations, and appropriate attribution language.
- 07Learning rhythm: owners review evidence on schedule, record adaptations and dissent, and update the program model and next question.
Primary references
Sources and review
- Coach House Accelerator
Coach House
The internal sequence behind this guide: need, theory of change, systems thinking, program piloting, evaluation, and explicit assumptions.
- CDC Program Evaluation Framework
Centers for Disease Control and Prevention
Current six-step federal framework with collaborative engagement, fair and just practice, use of insights, and five evaluation standards.
- Step 3 — Focus the Evaluation Questions and Design
Centers for Disease Control and Prevention
Guidance for connecting purpose, intended users and uses, questions, program stage, feasibility, context, and evaluation design.
- Step 4 — Gather Credible Evidence
Centers for Disease Control and Prevention
Guidance for selecting methods, indicators, measures, sources, quantity, quality, timing, and context appropriate to the question.
- Step 5 — Generate and Support Conclusions
Centers for Disease Control and Prevention
Guidance for analysis, interpretation, recommendations, and conclusions supported by evidence and context.
- Step 6 — Act on Findings
Centers for Disease Control and Prevention
Guidance for planning use, preparing findings, and facilitating insights into action with intended users.
- Evidence Readiness Resources
AmeriCorps
Federal training and templates for logic models, research questions, evaluation planning, data collection, reporting, and use.
- Protecting Personal Information: A Guide for Business
Federal Trade Commission
Federal guidance to know what personal information is held, keep only what is needed, protect it, dispose of it securely, and plan for incidents.
This educational guide and planning tool do not determine program effectiveness, causal attribution, evaluation quality, grant compliance, research status, participant consent, privacy compliance, or impact. Requirements and appropriate methods vary by question, population, program, funder, jurisdiction, risk, and intended claim. Review consequential plans and findings with affected people and qualified evaluation, research, privacy, legal, accessibility, data-security, and subject-matter professionals as appropriate.