MavenlyMavenly
    Field Notes

    Practice · 12 min read

    You probably have more outcome data than you think — it is just not in a usable form.

    Small organizations rarely lack data. They lack a collection habit, a definition of what counts, and a place where the numbers survive staff turnover.

    The Mavenly Team·Mavenly Practice Team·July 17, 2026

    Reviewers ask three evidentiary questions: is the need real, does your approach work, and have you delivered before. Organizations that cannot answer these lose to organizations that can, and the gap is almost never about program quality. It is about whether anyone wrote things down in a form that survives.

    The first thing worth establishing is that most small nonprofits are not data-poor. They are data-disorganized. Attendance is in sign-in sheets. Outcomes are in a case manager's notes. A university partner ran an evaluation in 2022 that nobody has cited since. Participant feedback exists as a folder of paper surveys. The information is real; it is simply not in a state where anyone can produce a number under deadline.

    So the first project is not a data system. It is an inventory: what do we currently collect, where does it live, who owns it, and what could we report today if asked. That exercise alone typically surfaces two or three usable metrics an organization had forgotten it had.

    For demonstrating need, the standard is a mix of public data and your own. Public sources — Census and American Community Survey, county health rankings, state education dashboards, BLS — establish the condition credibly and cost nothing. Your own data establishes that the condition is present in the specific population you serve, which is the part reviewers actually weigh. A waitlist count is evidence. A referral log is evidence. Demand you turned away is among the most persuasive numbers available and almost nobody records it.

    For demonstrating that your approach works, the honest hierarchy runs from external evaluation to pre/post measurement to output counts, and there is no shame in being where you are as long as you say where you are. An organization reporting 'we served 340 participants and 71 percent completed the program' is credible. An organization implying outcomes it did not measure is not, and reviewers detect the difference easily. Claiming a 'documented 40 percent improvement' without naming the instrument invites a question you cannot answer.

    The practical build for a small organization is smaller than the term 'evaluation capacity' suggests. Pick three to five indicators, not twenty. Define each one in writing so that two different staff members counting the same thing get the same number — 'participant' means what, exactly; 'completed' means what threshold. Record the definition, the collection method, the frequency, and the owner. Collect at the moment of service rather than reconstructing later. And keep a baseline, because change is unreportable without one.

    Definitional discipline is the part that is skipped and the part that causes the most damage. Two program sites counting participants differently produces a number the finance director quietly distrusts, and that distrust eventually reaches a funder report. Write the definitions down once.

    Where existing capacity can be borrowed, borrow it. University partnerships are frequently free — graduate programs in public health, social work, education, and public policy need field placements and capstone projects. AmeriCorps VISTA members can build data systems. Funders themselves increasingly pay for evaluation capacity; several will fund it as a line item if you ask, and asking is a stronger move than pretending you already have it.

    On analysis, the bar for grant purposes is far lower than academic work. Counts, percentages, changes over time, and comparison to a relevant benchmark cover nearly everything a reviewer wants. Sophistication is not rewarded; clarity and honesty about method are.

    It is worth saying explicitly what should never be done, because the temptation is real at 1 a.m. with a section unfilled: do not estimate a number and present it as measured. Funders remember reported figures, they appear in your final report obligations, and an outcome you invented becomes a compliance problem the moment the grant is awarded. A flagged gap costs you points once. A fabricated number costs a relationship.

    Past performance is the third question and the easiest to prepare for, because it is purely archival. For every completed grant: the award amount, the period, what was promised, what was delivered, and what you learned. Keep it as a running record. Organizations that maintain it can answer prior-performance sections in twenty minutes; organizations that do not reconstruct them from memory and understate their own results.

    The compounding argument is the one to hold onto. Data collected this year is what makes next year's proposal competitive, and the year after that, and it is also what makes the final report writable without a scramble. The organizations that look data-rich in year three are almost never the ones that bought a system. They are the ones that defined five indicators in year one and never stopped collecting them.