There is a category of work that people analytics teams do that never appears on a budget line, isn't tracked as a deliverable, and is rarely discussed in team retrospectives. It is the work of reading open-ended survey responses and turning them into something a leadership team can actually use. This work happens before the insight document exists. It is invisible until it's done, and when it's done poorly, no one outside the team can tell. That combination of invisibility and consequence is what makes manual summarization one of the most expensive practices in HR analytics, and one of the hardest to challenge.
What Manual Summarization Actually Involves
When a pulse survey closes, the quantitative results go into a dashboard almost immediately. Score distributions, eNPS, trend lines: all of that is computed. The open-ended responses, which often contain the most contextually specific signal in the whole survey, require a different process.
Someone, usually an analyst or a more senior HR professional, downloads the text data. They read through it. They take notes or highlight as they go. They begin to see patterns. They write a summary, typically in a PowerPoint slide or a short document, that characterizes what themes appeared most often. That summary goes into the report deck. Leadership reads the summary. The raw responses do not.
For a survey with 300 to 400 open-ended responses, this process takes somewhere between half a day and a full day of focused reading, depending on response length and analyst experience. For a survey with 2,000 responses, it can take a week, especially if the analyst is trying to be thorough. For 4,000 responses, most teams either extend the timeline significantly or stop being thorough. Both options are costly.
The Time Cost Is Only Part of the Problem
If manual summarization were only slow, the solution would be straightforward: add analyst time. The deeper problem is that the summarization process introduces systematic bias that doesn't announce itself in the output.
Human readers of text data tend to cluster toward the familiar. When an analyst reads 400 open-ended responses, they are scanning for patterns, but their pattern recognition is shaped by what they already know about the organization. Issues they've heard mentioned before in meetings register quickly. Issues they haven't heard about before take longer to see as patterns. A theme that appears in 30 responses, all expressed in slightly different language, may not feel like a theme until the analyst has read far enough into the dataset. If they stop at 200 responses, they may not have seen enough instances of it yet.
There is also what might be called the recency effect in reading. The last 50 responses an analyst reads tend to be more present in their mind when they sit down to write the summary. This is not a failure of individual analysts; it is a predictable property of how human short-term memory works. The result is that the content of the summary can be influenced by the order in which responses were processed, which is typically the order in which they were submitted. Early respondents, who may differ systematically from late respondents in engagement level, department, or seniority, can be underweighted simply because their responses were read earliest.
The Sampling Problem That Nobody Talks About
When analysts face 2,000 or more responses and have limited time, they sample. This is almost never stated explicitly. The report says "themes from employee feedback," not "themes from the first 400 responses we had time to read." But the summary reflects a sample, and the sample has selection properties that are almost never characterized.
Suppose an analyst reads the first 500 responses from a 2,500-response survey and then skims the rest. If the survey platform presents responses in submission order, and employees in one business unit tended to submit earlier than employees in another, the analyst's sample is demographically skewed in a specific direction. The summary then reflects that skew without any annotation that this skew exists. Leadership reads the summary without knowing it represents one segment of the organization more than another.
This is not a hypothetical problem. It is the default problem whenever human reading capacity is the binding constraint on how much of a dataset gets analyzed. The only complete defense against it is analyzing the full dataset, not a subset, which manual methods cannot reliably do at scale.
Calibration Drift Across Reviewers and Cycles
In teams where multiple analysts contribute to the summarization process, a different problem emerges. Different analysts apply different implicit thresholds for what counts as a "theme." One analyst might require 20 or more mentions of a concern before calling it a theme in the summary. Another might flag something at 8 mentions if the language is particularly vivid or urgent. Neither standard is formally wrong, but they produce different summaries from the same data.
This calibration drift also happens across time, not just across people. An analyst who summarized pulse results six months ago has a different reference frame than the same analyst doing it now, partly because their expectations have been shaped by the most recent organizational events. A period of organizational tension causes analysts to weight conflict-related themes more heavily than they would in a calmer period. Whether that is appropriate or distorting depends on the specific situation, but it should be a conscious choice rather than an artifact of the reviewer's current state of mind.
Formal intercoder reliability processes, borrowed from qualitative research methodology, can address this. But applying rigorous intercoder reliability to every pulse survey cycle at scale is impractical for most HR teams. The overhead required to develop a coding scheme, train multiple reviewers, and measure agreement percentages would take longer than just reading the responses twice.
What the Summary Does Not Preserve
A manual summary, well-executed, preserves the major themes and their relative frequency. What it does not preserve is the texture of the data: the specific language employees used, the combinations of concerns that appeared in the same response, the intensity of feeling that distinguished some responses from others, and the contextual detail that reveals why an issue is arising at this particular time rather than three months ago or three months from now.
Leadership teams working from summaries are working from an abstraction of an abstraction. They have the analyst's interpretation of what was present in the data, not the data itself. When they ask follow-up questions, they are asking about the summary, not about the employee population. "Were people specifically worried about X or was it more about Y?" is a question the summary cannot answer, but the response corpus can.
This matters most in situations where the decision at stake is significant: a structural reorganization, a return-to-office policy change, a benefits revision. These are precisely the moments when HR leadership most needs granular fidelity in the feedback they are working from, and precisely the moments when the manual summarization process is most likely to be under time pressure and therefore least likely to be thorough.
The Organizational Capacity Argument
For growing organizations, the manual summarization bottleneck compounds in a specific way. As headcount increases, survey response volume increases proportionally or faster, if engagement programs are expanded at the same time. But analyst headcount tends to be set once, or grows much more slowly than the organization. A team that managed pulse surveys adequately at 500 employees is managing the same survey volume a year later at 1,200 employees, with either the same staff or one additional hire.
The practical response is typically to reduce the frequency of analysis cycles, to shorten reading time per cycle, or both. Each of those choices reduces the fidelity of the feedback loop without anyone formally deciding to reduce it. The organization is nominally running the same listening program, but the quality of the analysis is lower. The gap between what employees said and what leadership hears has quietly widened.
Addressing this does not require believing that automated analysis is always better than human analysis. It requires recognizing that manual analysis has a hard capacity ceiling, and that ceiling is often reached well before the organization becomes large enough for the problem to be visible on a budget line. The cost of manual summarization is paid in analytical coverage, in bias, and in organizational decision quality. Most of it never shows up as a line item anywhere.