Existing data becomes human-subjects research when you obtain, use or analyse identifiable private information about living people for a systematic investigation designed to produce generalizable knowledge. Identifiability decides it, not the source of the file. You file for that reading; you never declare it.
Does working from records count as research at all?
Two federal definitions run in sequence, and skipping the first is where most records files go wrong. Research, at 45 CFR 46.102(l), is a systematic investigation designed to "contribute to generalizable knowledge". The second definition covers the people: a living person from whom an investigator gathers information by interacting, or about whom the investigator "studies, analyzes, or generates identifiable private information" (46.102(e)(1)).
Read those in order, because the reviewer does. If the aim is honestly to see whether last quarter's fall rate moved on your unit, and to tell your unit about it, the activity may not meet the research definition at all, and the rest never applies. If the answer is meant to travel beyond the site, you are inside that definition, and only then does identifiability decide anything. OHRP sets out the same sequence: research, then human subjects, then exemption.
The trap is that one chart pull can sit on either side. A run chart posted on the unit board and a retrospective comparison written for an audience beyond the unit can draw on identical rows. What separates them is the intent you wrote down, and intent stated in one document has to survive in all the others. The reading we issue starts there for that reason.
What actually makes a dataset identifiable?
The regulation at 46.102(e)(5) is unusually plain. Private information is identifiable when the subject's identity "may readily be ascertained" by you, or is "associated with the information". Two phrases, two different tests. The first asks not whether you intend to look, but whether looking is available to you. The second means a name sitting beside a row counts even if you never read that column.
This is why a retained key is the most consequential object in a records project. Replace names with P01 and P02, keep a linking file in a folder only you can open, and the information stays identifiable to you, because you can resolve any code whenever you choose. De-identified describes what exists, not where it is filed.
The reviewer is not asking whether you would look. The reviewer is asking whether you could.
Where does a coded extract land?
OHRP's 2008 guidance on coded private information answers this, and it is the document reviewers reach for. Work involving only coded information falls outside the human-subjects definition when two conditions hold together. The information was not gathered for this particular project by interacting with anyone, and the investigator "cannot readily ascertain the identity" behind the codes, because a written agreement bars release of the key, or a repository's approved policies bar it, or the law does.
Notice the practical cost. Somebody else must hold that key, and the arrangement has to be written down before the extract moves. A friendly understanding with the analyst who ran the query is not what the guidance describes. That guidance predates the 2018 revision and cites older section numbers, so read its reasoning against the current text at 46.102(e) and 46.104(d)(4).
Then the closing recommendation, which surprises people most: OHRP advises that investigators should "not be given the authority" to decide on their own that coded-data work involves no human subjects. WGU's policy page says the equivalent, requiring proposals to reach the IRB so the IRB establishes the level of review. The reading is the office's to issue.
| What you plan to obtain | What the reviewer looks at | Usual reading |
|---|---|---|
| An aggregate report from the site: counts and rates only | Whether any cell is small enough to name a person | Frequently not human-subjects research |
| A row-level extract you pull yourself, carrying record numbers or service dates | Identifiability at the moment of recording | Human subjects; expect a review path |
| A coded extract, key held by the site's analyst under a written agreement not to release it | The agreement itself, and when it was made | Often outside the human-subjects definition |
| A coded extract, key retained by you "in case of questions" | Your ability to re-identify at will | Identifiable; the key decides it |
| A public dataset published for open use | That it is genuinely public, not merely obtainable | Named in the secondary-research category |
| Data held by WGU about its own population | The institutional permission that precedes the IRB | Separate approval first; see below |