Blog

Follow the data

Ask an organization where its sensitive data lives and the answer usually names a system. Ask where it goes and the answer takes considerably longer, because the honest version involves an export somebody built years ago, a spreadsheet that leaves by email once a month, and a vendor portal nobody in IT provisioned.

That second answer is the one that matters, though not quite for the reason usually given. Two in-scope systems talking over an in-scope network leak nothing: every point on that path already meets the requirements, and the data never leaves the protected environment. The trouble starts where a flow crosses a boundary—into a system nobody assessed, onto a network that is not in scope, out to a provider, or into somebody's mail.

Internal flows still have to be written down. A desktop reading from a file server across an in-scope link is not a risk, but an organization that cannot describe that path cannot show the link is in scope, cannot tell when the path changes, and has no way to notice the day somebody adds a hop that does leave.

What a flow has to capture

A useful data flow follows a category of data through its whole life rather than mapping the network. That presumes the categories exist: an organization that has not decided what counts as sensitive data will draw one enormous flow containing everything, which answers nothing.

In practice the unit is the project rather than the system, because that is how data arrives and leaves. Project A receives data from sources B and C, processes it on systems D and E, and returns a result to the customer by method X. Project F has a different shape even when it runs on the same infrastructure. One flow per project, per category of sensitive data, is the granularity that stays accurate. For each kind of sensitive data:

The copies are where this exercise earns its cost. Most organizations can name the primary system immediately and have never enumerated the derivatives.

What happens when the project ends belongs in the flow too, and it is the part most often missing. Sometimes the data is destroyed. Sometimes it has to be kept for a stated number of years because a contract or a regulation says so, which means somebody is storing and protecting it long after the work that produced it finished, and somebody is paying for that. Either way the obligation outlives the project. The flow is where it should be recorded—along with whether the project budgeted for it, which is commonly forgotten and is discovered at the point the cost has already been incurred.

The flow determines the threat

The reason this belongs in a risk assessment rather than in a diagramming tool is that the route changes what can go wrong.

Consider two companies moving identical data between two sites. The first writes it to write-once optical media and sends it by courier. Its risks are physical: the package goes astray, the custody record is incomplete, and the media cannot be revoked once it has left the building, because there is no remote wipe for a disc. Its mitigations are correspondingly physical—encrypt the media, log custody, control who can produce a copy.

The second uploads the data to a cloud service and lets the far site download it. Its risks are entirely different: an account that can be phished, an administrator at the cloud service provider who is outside the organization's control, a sharing setting that quietly permits more than intended, and a copy that persists on someone else's hardware after the contract ends.

Neither arrangement is wrong, and neither is safe in the way the other is. An organization that has not documented which one it is using cannot say which set of risks it has chosen, and a generic control list will not tell it.

That comparison is the argument for doing the work at all. The flow is the input a risk assessment needs in order to identify the risks that actually exist rather than the ones a template suggests. But it does something else too: it is how an organization establishes that every point under its control is doing the right thing with the data.

Protection has to travel with the data, and the obligation does not stop at the edge of the primary system. In the first company, the courier needs terms about custody and loss, and the media needs encrypting—because the control that matters is the one accompanying the disc, not the one on the server it was written from. In the second, the cloud service needs the authentication the policy requires, a sharing configuration somebody has actually checked rather than assumed, and a contractual answer about deletion when the term ends.

Every hop is a place where the organization either does the right thing or finds out later that it did not. Enumerating the hops is what makes that checkable, because a control cannot be applied to a path nobody has written down.

Standards ask for this, in nearly the same words

The requirement recurs across frameworks.

NIST SP 800-171 expects the system security plan to describe the system boundary, the operating environment, and how controlled unclassified information moves through and beyond it. In a CMMC assessment this is inseparable from scoping: the flow is what establishes which systems are in scope, and an assessor with an unconvincing flow diagram will widen the boundary rather than narrow it.

PCI DSS is the most explicit of the group. Version 4.0 requirement 1.2.4 calls for an accurate data-flow diagram showing all account data flows across systems and networks, kept updated as the environment changes. Version 3.2.1 numbered the same obligation 1.1.3, which is worth knowing before quoting a number from memory.

CIS Controls safeguard 3.8 asks for documented data flows including those of service providers, reviewed annually or when significant changes occur.

NIST Cybersecurity Framework ID.AM-03 asks that representations of internal and external network data flows be maintained, including flows to third parties and to infrastructure services.

The convergence is not a coincidence. Each of these frameworks has to define what it applies to, and the honest way to do that is to follow the data.

Why it gets skipped, and what that costs

Documenting flows is unglamorous, and the resistance to it is predictable. In practice it arrives most often from project managers, who are measured on delivery dates and see the exercise as a tax on a schedule that is already tight. The data is moving perfectly well without a diagram. The diagram can be written afterward, when there is time.

There is rarely time afterward, and the argument misjudges what is being deferred. Three things follow from skipping it.

The first is the obvious one. Undocumented data movement is unprotected data movement, because nobody applies controls to a path nobody has described.

The second is that the omission is itself a control failure. The standards above are requirements rather than suggestions, and for a defense contractor the requirement has a contract behind it: implementing NIST SP 800-171 is mandated by DFARS 252.204-7012, a clause the company signed. An organization that has not documented its flows is not merely disorganized. It is failing a requirement it agreed to in writing.

The third follows from the second, and is the one worth raising when a schedule argument is being lost. Failing to meet a contractual security requirement is a breach of that contract. In the federal context the exposure does not stop there: an organization that has affirmed a compliance it does not have—and CMMC requires such an affirmation every year—can face liability under the False Claims Act, which the Department of Justice has pursued specifically against cybersecurity misrepresentation.

That is a considerable distance from a conversation about whether a diagram is worth two days, which is rather the point.

It is worth being honest about that cost, because the objection usually inflates it. Only a genuinely complicated environment takes days. A simple flow can be written in a couple of sentences: the data arrives through DoD SAFE, is stored on server abc, processed on compute server def, visualized on desktop ghi, and returned to the customer through DoD SAFE again; backups go to backup servers jkl and mno; every network between those systems is in scope. That is a complete flow for that project. It takes about a minute to write, and it is enough to reason about. The diagram that takes two days is usually the one nobody broke into projects first.

The flow includes people

A data flow that names only systems is half a flow. Data is not moved solely by integrations; it is read by people, and who can see it at each stop belongs in the description.

For most organizations that list is expressed as group membership. Access is granted by role, the role maps to a group, and the group is what actually determines who can open the file. That is the right design, and it has a well-known decay: groups accumulate. Someone joins a project and is added. The project ends. The person transfers to another department, and nothing in the process that moved them removes them from the group—often because the person who granted it has themselves moved on.

The result is a flow that is accurate about systems and quietly wrong about people. The data goes exactly where the diagram says it goes, and is readable by rather more people than the organization believes.

The remedy is a periodic review of both, and the standards ask for it. NIST SP 800-171 requires that the privileges assigned to roles or classes of users be reviewed at an organization-defined frequency to validate that those privileges are still needed, and for defense contractors the department has set that frequency at least every 12 months.

Two things make the review worth conducting. Compare membership against a current statement of who should have access, rather than against last year's list, because comparing a list to itself confirms nothing. And have the business owner do the confirming rather than IT, since IT can say who has access but not who needs it.

Both reviews—of the flow and of the group membership—have to leave a record. An assessor asking whether access is controlled is not answered by a current screenshot of a group, which describes only today. What answers it is a series of dated reviews showing that someone has been checking, what they found, and what they removed. A history of protecting the data is a different claim from a snapshot of it, and only one of them is evidence.

Keeping it honest

A flow diagram drawn once and filed is worse than none, because it invites confidence in a description that has stopped being true. Three habits keep it honest.

Trigger the review from events that already happen. A calendar review alone will always lag, because whatever invalidated the diagram happened in March and the review is in November. The events worth attaching it to are the ones the organization already has a process for: a new vendor or service provider, a new integration between systems, a department changing how it works, an acquisition, a new product line, a system replaced or retired, a contract ending, and any real change in where people work from.

The practical move is to make the flow review a step inside a process that already runs—change control, procurement approval, vendor onboarding—rather than a separate obligation somebody has to remember. Better still if those are not three processes: where procurement and vendor changes are themselves raised through change control, there is one place a change to the environment is recorded and one place the flow review hangs from. An obligation that depends on memory depends on one particular person still working there.

Validate against observation rather than recollection. The diagram records what the organization believes. Several systems already record what actually happened, and each catches a different kind of omission:

None of these is complete alone. Together they are usually enough to find whatever the diagram is missing.

An organization already running data loss prevention has one more: tagged files can be counted, so the volume moving along each path becomes something measured rather than estimated. That is a real capability and an expensive one—to buy, and more so in the human time good tagging takes, since the shortcuts that make tagging quick are the ones that make it wrong. It also tends to commit an organization to a single vendor from the server through the desktop to the network infrastructure. Worth using where it exists; not worth buying in order to draw a diagram.

Keep it readable, and keep it dated. A diagram that has grown to cover every system at maximum detail stops being maintained, because updating it becomes a project. One flow per category of sensitive data, drawn at a granularity where a reader can actually follow the path, is more useful and survives longer. Record the date of each review, who performed it, and what it was checked against—both because next year's reviewer needs that and because it is the evidence the review happened at all.

The gap between the documented flow and the observed one is the finding. It is also, reliably, where the unregistered vendor and the undocumented export turn up.

Kenneth Ingham Consulting helps organizations map their sensitive data flows as part of scoping and risk work, which is usually where an organization discovers how far its data actually travels. It is not work that can be done from the outside: the people who know where the data really goes are the ones who handle it, and the map is the product of a structured conversation with them.

Tagged: data flows, scoping, risk assessment, CUI