PII Data Inventory From Code: Know Your Regulated Data
Ask a security team what personal data their software handles and you will usually get a spreadsheet. Someone ran a survey, application owners filled in a form, and the answers were accurate on the day they were written. A PII data inventory built that way starts going stale the next time anyone merges a pull request.
The code already knows the answer. Every model, table and schema your application persists is declared somewhere in source. Every API contract you publish says what goes in and what comes out. A data inventory can be derived from that, kept current on every analysis, and handed to an auditor as a report instead of a questionnaire.
This post explains how a code-derived data inventory works, what it can and cannot tell you, and how to use it when privacy, audit or incident response come asking.
Why a PII data inventory matters to security teams
Data inventories used to be a privacy office concern. They are now a security input, for three practical reasons:
- Regulations require one. GDPR Article 30 obliges controllers to maintain records of processing activities, including the categories of personal data involved. See the text of Article 30. Other regimes ask similar questions in different words.
- Risk depends on what the data is. A SQL injection in a service that stores email addresses is a different incident from the same bug in a service that stores payment details or health records. Prioritization without data context is guesswork.
- Incident response needs it fast. When a service is breached, the first question from legal is "what data was in it?". An inventory that takes a week to assemble is not an answer.
The common thread: the inventory has to be current and it has to be tied to the systems that actually hold the data.
Surveys vs code: where each inventory comes from
A survey-based inventory records what people believe their systems hold. A code-derived inventory records what the systems declare. The difference shows up in predictable places:
- New fields. A developer adds a
date_of_birthcolumn in a sprint. The survey does not know until next year. The code knows on the next analysis. - Forgotten services. The internal tool that copies customer records into a reporting database was never on anyone's list. Its models are still in a repository.
- API exposure. A survey rarely asks which endpoints return personal data. An API contract says so directly.
- Ownership. A code-derived entry comes with a repository, a module and the team that owns it.
Surveys still have a role. Code cannot tell you the lawful basis for processing, retention agreements with vendors or whether a field is populated in production. Those are process facts. Use code for the "what and where", and people for the "why and under what terms".
How a code-derived data inventory works
The mechanics are straightforward to describe, even if the classification work behind them is not.
1. Extract persisted entities
Start from the declarations the code already contains. In practice that means the ORM models and schema definitions your frameworks use: SQLAlchemy and Django models in Python, TypeORM, Prisma and Sequelize in TypeScript and JavaScript, JPA entities in Java, GORM in Go, Entity Framework Core in .NET. Each model gives you an entity, its backing table or collection, and its fields.
2. Read the interface contracts
Entities tell you what the code stores. Contracts tell you what it moves. OpenAPI, GraphQL SDL, protobuf and JSON Schema definitions describe request and response bodies field by field.
3. Classify fields against a taxonomy
Each field is matched against a data taxonomy. A match produces a classification with:
- an attribute, the specific kind of data, such as an email address or a government identifier;
- a category, such as contact information or individual identifiers;
- a sensitivity, such as PII or financial data;
- a severity, rolled up to the entity and the repository as a ceiling;
- the regulations that treat the attribute as in scope. One attribute commonly maps to several.
Fields that match nothing stay unclassified but are still counted. That ratio is useful on its own: "4 of 12 fields classified" tells you how much of an entity is regulated.
4. Map regulations back to code
Once classifications carry regulations, you can turn the inventory around. Start from a regulation and list every attribute, entity and repository that brings it into scope. That is the question an auditor actually asks.
Consumes vs transmits: reading endpoint exposure
For APIs, direction matters. An endpoint that consumes personal data accepts it in the request. An endpoint that transmits personal data returns it in the response. They are different problems:
- Consumption is a collection and retention question. Why do we take this, and how long do we keep it?
- Transmission is an exposure question. Who can call this, and is it authenticated?
The strongest signal comes from putting data classification next to security posture. An endpoint that transmits critical personal data, is reachable from the internet and has no authentication is a far more urgent problem than the same payload behind an auth guard. Those two facts are usually in two different tools. They belong on one screen.
What a data inventory is not
Be precise about the limits, because auditors will be:
- It is not a compliance determination. An inventory says which data your code handles and which regulations map to it. Whether you comply depends on controls, contracts and processing purposes outside the code.
- It is not a production data scan. A declared field may be empty in practice. A code-derived inventory reports what the software can hold, which is the right starting point for risk but not proof of volume.
- It is only as good as its taxonomy. Classification by field name and structure catches the common cases well. Custom naming conventions and free-text blobs need review.
Using a code-derived PII inventory in practice
A few workflows get most of the value:
- Audit scoping. Filter to one regulation, switch to repositories, and hand the list with owning teams attached to the auditor. Filter endpoints by the same regulation to get the in-scope API list.
- Finding prioritization. Before deciding how to treat a finding, check what the affected service holds. Weight fixes on services with high-severity data.
- Pull request review. A new model in a pull request appears in the inventory once the code is analyzed, so privacy review can happen when a field is introduced rather than a year later.
- Incident response. When a service is implicated, pull its data tab first. That is the scope of the notification conversation.
- Financial controls. Not every regulation is about people. Sarbanes-Oxley scoping is about financial records such as payments, payroll and ledgers. Filter by sensitivity to isolate that footprint.
How Heeler builds the Data Inventory
Heeler's Data Inventory is built from your code on every repository analysis. It extracts the entities your code persists from the ORM models and schemas listed above, plus Rust row structs from Diesel, SeaORM and SQLx, and reads interface schemas from OpenAPI, GraphQL SDL, protobuf and JSON Schema. Each field is classified against a taxonomy Heeler maintains, with attribute, category, sensitivity, severity and regulations. Health data is flagged separately rather than inferred from its category. Nothing is tagged by hand.
The inventory can be browsed six ways: by entity, attribute, category, regulation, repository and endpoint. The regulations view covers frameworks including GDPR, CCPA, HIPAA, India's DPDP Act, ISO/IEC 27001, US state privacy laws and Sarbanes-Oxley. The Endpoints view marks each API endpoint as consuming or transmitting personal data, and the endpoint detail puts data classification beside its security posture: whether it is public and whether anything protects it.
Repositories, services and applications each carry a Data tab scoped to that asset, and a Compliance Report produces a point-in-time document for the current scope, filtered by environment and match confidence so sandbox repositories and low-confidence matches stay out of reported coverage. The inventory sits in the same catalog as your services, endpoints and findings in the Context Engine, so the data a service holds is one click away when you set its tier or decide how to treat a finding against it. See the Data Inventory docs and Compliance and Audit-Readiness for more.
A checklist for your first code-derived inventory
- Connect every source control organization, not just the flagship ones. The forgotten services are the point.
- Pick one regulation your auditors care about most and review its in-scope repositories first.
- Review the unclassified-field ratio on your highest-tier services. Low classification on a sensitive service usually means naming conventions the taxonomy missed.
- List every endpoint that transmits high-severity data and check its authentication.
- Assign each in-scope repository an owner who confirms the "why" that code cannot supply.
- Regenerate the report before each audit rather than updating last year's.
FAQ
What is a PII data inventory?
A PII data inventory is a record of the personal data an organization's systems hold, where it lives and which regulations apply to it. It is required or expected under GDPR and many other privacy regimes.
Can you build a data inventory from source code?
Yes. ORM models, schema definitions and API contracts declare the entities and fields your software stores and moves, and those fields can be classified against a data taxonomy on every analysis.
Does a code-derived inventory replace a privacy assessment?
No. It tells you what data the code handles and where. Lawful basis, retention and vendor terms still need a human process.
What is the difference between an endpoint consuming and transmitting PII?
Consuming means personal data arrives in the request. Transmitting means it leaves in the response. Transmission is usually the more urgent exposure question.
How often should a data inventory be updated?
Every time the code changes. A code-derived inventory updates on each analysis, so a new field appears once the pull request that adds it is analyzed.
See your data inventory
If your answer to "what personal data do we hold?" is still a spreadsheet, see what your code says instead. Heeler builds the inventory from your repositories and ties every entry to the code, the endpoint and the team that owns it. Get a demo.


.jpg)
