Skip to main content

Equity in AI for Public Health: Part 2 – From Clusters to Careful Personas

·9 mins

Plain-language takeaway: Clustering can help public health teams see how needs, strengths, risks, and supports overlap. Equity changes how the data is prepared, how patterns are tested, and whether a pattern should become a persona.

Part 1 explained why equity in AI must guide public health work. It also introduced the four actions in NIST AI RMF 1.0: govern, map, measure, and manage.

The second article illustrates what these commitments look like in practice. It follows the WHY Survey Personas project from survey preparation to human review. The examples given below are drawn from the project’s workflow, but no sensitive findings, subgroup differences, cluster sizes, or draft persona labels are disclosed.

What exactly are clusters and personas? #

Clustering is a method that compares many responses and places more similar response patterns closer together.

Imagine sorting a mixed group of buttons by size, shape, colour, and number of holes. Different sorting rules create different groups. Clustering works similarly, but survey responses are more complex, and the choices have bigger impacts.

While the method can identify similarities, it cannot determine whether a group is important for public health. It cannot explain why a pattern exists or prove that one thing has caused another. Furthermore, it cannot demonstrate that all the students in a group have the same lifestyle or needs.

In this project, a persona is a plain-language summary of a general response pattern. It does not represent a real student, a diagnosis, a prediction or a fixed kind of person.

A cluster is a group created by a computer. A persona is a carefully reviewed explanation of that group. Not every cluster should become a persona.

The gap between these two steps—computer output and public health meaning—is where much of the equity work takes place.

How equity changes the WHY Survey workflow #

The WHY project does not rely on just one score or chart. It uses a series of checks that connect the technical work with public health judgment.

1. Start with the purpose and prohibited uses #

Before grouping any responses, the team must be clear about what the work is for.

The intended purpose is exploratory: to help WDGPH staff understand connected patterns and then formulate better questions regarding the needs and the support required at a population level.

The project is not intended to identify or classify students, predict individual outcomes, or rank schools or communities. It cannot determine who receives a service, support, discipline, or surveillance. Nor can it replace professional and community judgment.

These limits are not just a note added at the end. They affect what information is used, what outputs are created, where detailed files are stored, and what a reviewer can approve.

2. Respect differences in the survey #

The questionnaire used in the WHY Survey is not the same for all students, since different questions are given to junior, intermediate, and senior students and some questions only appear after a certain earlier answer has been given.

If the results from all the students were combined together, the differences might cause the outcome to be distorted. The computer could end up grouping the students according to the survey form they got, rather than according to relevant health patterns.

So, the project uses division-specific analysis and other methods that consider survey routing. In simple terms, students are compared within the right survey groups before making larger comparisons.

This is a clear Equity in AI decision. The survey’s structure is seen as part of the evidence.

3. Ask why an answer is “Not Asked” #

“Not Asked” does not always mean the same thing.

Consider two examples:

  • A question may be deliberately omitted from the junior survey. In that case, the response should not be altered to say “No.”
  • A follow-up question may appear only when a student answers “Yes” to an earlier question. For instance, a detailed question about sports betting will only appear after the student has reported having engaged in sports betting. If an earlier answer makes the follow-up behaviour impossible, the project can create a separate value for a particular analysis. The initial reply remains “Not Asked.”

The same care is still needed in this second situation. If a student states that they had not been bullied, then a question concerning whether an adult’s response was helpful does not apply. It would be incorrect to say “No, the adult was not helpful.”

The original values are therefore maintained. New analysis values are recorded in a separate, documented layer and structural “Not Asked”, normal missingness, and “Prefer not to answer” are all kept apart.

Why does this matter? Because treating every blank the same could make a survey rule look like student behaviour. Even a small cleaning decision could shape the clusters and create an unfair story about youth.

4. Give variables different jobs #

Not every variable needs to play the same role.

Some measures of health, well-being, risks, and supports may be used to explore possible clusters. Sensitive equity and demographic information may instead be kept out of the persona-defining features and used later in a separate fairness review.

For example, after a possible clustering solution is created, approved reviewers may ask whether some groups are represented differently among the proposed personas. That difference is a signal for investigation, not proof that the method is biased or permission to define a persona by identity.

Several explanations may still be possible, including survey participation, question wording, missing information, social conditions, the clustering method, or chance. The project must therefore record the concern and decide what further review is needed.

This separation helps avoid making personas based on identity, but still lets the team check if the method may represent some youth differently.

5. Compare reasonable alternatives #

A computer can produce clusters even when the groups have little public health value. The project therefore compares different ways of preparing the information and finding patterns.

The review asks:

  • Does a general pattern reappear when a reasonable method or approach is altered?
  • Is the pattern clear enough to explain without guessing?
  • Will it provide useful information for public health?
  • Might the survey routing or the missing information be creating it?
  • Could a simpler explanation account for it just as well?
  • Could it hide, misrepresent, or stigmatize some youth?

A cluster may look clear within one method but change when another reasonable method is used. Some students may also sit near the boundary between two patterns rather than fitting neatly into one.

The WHY project tracks this uncertainty. Terms like core, boundary, and method-sensitive are used to describe confidence in the method, not to classify students. This supports a softer view: personas are overlapping summaries, not strict boxes that all young people fit into.

No single technical score determines the outcome; a candidate may be advanced for further review, restricted, revised, kept merely as supporting evidence, or rejected.

6. Check equity, privacy, and confidentiality together #

Once candidate solutions have been identified, the project should carry out its own review concerning equity, privacy, confidentiality, and fairness using the approved aggregate information.

Reviewers should check if there are differences in terms of representation, missingness, or stability among the equity-related variables that have been approved. They also need to make clear who is included in each calculation. The groups consisting of all students, those who saw the question, those who answered it, and those included in an analysis are not always the same.

Small results are suppressed under approved privacy rules. To maintain confidentiality, row-level labels, unsuppressed counts, and detailed subgroup joins remain accessible only to authorized staff working within approved private systems. Public materials use only reviewed summaries.

Different persona proportions across groups do not, by themselves, prove discrimination, causation, or an unfair method. Similar proportions do not prove fairness. Hiding a result does not mean there is no concern.

The audit is therefore a diagnostic and governance tool, not a fairness certificate.

7. Turn evidence into respectful language #

A candidate cluster should become a draft persona only after its pattern, uncertainty, privacy, and possible uses have been reviewed.

At this stage, wording matters. “Students in this response pattern reported…” is safer and more accurate than “These students are…”. The first describes group survey information, while the second can sound like a fixed judgment about identity.

Draft persona descriptions should:

  • describe a pattern rather than a type of person;
  • include strengths and supports as well as needs;
  • separate cluster-defining information from later validation context;
  • avoid diagnosis, blame, moral judgment, and identity-based names;
  • explain uncertainty and what the persona cannot tell us; and
  • remain provisional until reviewed by people with different types of knowledge.

This review may include technical, epidemiology, privacy, health-promotion, community, and youth perspectives.

If the evidence does not support a respectful and useful explanation, the team should revise, limit, pause, or reject the persona. Creating fewer personas, or even none, is a valid result.

A practical example from beginning to end #

Imagine a team is reviewing a prospective response pattern involving several areas of well-being and support. The following example describes the process, not an actual project finding.

First, analysts confirm that the responses came from comparable survey forms. Next, they check whether “Not Asked” values reflect survey routing or actual missing answers. Then, they compare more than one clustering approach to see whether the pattern is reasonably dependable.

After that, approved equity information is used in a separate review to look for representation or measurement concerns. Small results remain suppressed. Public health and governance reviewers then ask whether the pattern has a useful planning purpose. They check that its description includes context and strengths. They also consider whether its name could stigmatize youth.

Only then can the cluster move forward as a draft persona. If any step raises a serious concern, the team goes back to an earlier decision or stops the candidate from moving forward.

This is how NIST AI RMF 1.0 looks in practice:

Framework actionWHY workflow example
GovernSet the purpose, review roles, privacy rules, and prohibited uses.
MapUnderstand survey routing, affected groups, public health context, and possible harms.
MeasureCompare candidate solutions and examine uncertainty, representation, missingness, and privacy.
ManageRevise, restrict, pause, reject, or cautiously advance a candidate based on the review.

From principles to accountable practice #

Part 1 explained why equity, privacy, and human responsibility must guide public health AI. Part 2 shows how these commitments change real decisions in the WHY Survey Personas project.

Together, the two articles make one argument:

AI can help public health teams see complex patterns, but responsible practice determines whether those patterns are trustworthy, respectful, and right to use.

For WDGPH, success is not just about finding clusters. Success means building a process that others can review. That process must respect survey design, protect youth, test uncertainty, look for unfair effects, and explain its limits. People are still responsible for every important decision.

If you want to apply these ideas to your own work, start by reviewing how you use data and AI tools with an equity, privacy, and confidentiality lens. Clearly define your project’s purpose, set boundaries for fair use, and involve colleagues from different backgrounds throughout the review process. Share questions, challenges, and lessons learned with your team. By taking these steps, you can help make sure every decision supports both public health goals and a commitment to equity.