Data Disclosure Avoidance Guidance

What is Data Disclosure Avoidance Guidance?

What is the purpose of this Data Disclosure Avoidance Guidance?

This document provides guidance for the production of data products (e.g., data tables, dashboards, reports, visualizations) containing student data so that individual identities are not inadvertently disclosed revealing personal information. 

To whom does it apply?

The guidance applies specifically to employees in the CSCU Office of Decision Support and Institutional Research (DSIR).  It also applies to entities who are authorized by DSIR and CSCU Academic Affairs through data sharing agreements to conduct analysis or evaluations with CSCU student education records.  Other institutional research professionals and CSCU employees who create and disseminate data products that include information about students may also follow this guidance as a representation of best practices.

When does it apply?

This guidance applies whenever data products containing information about students are produced for use by someone other than School Officials (see definition) as defined by the individuals’ institution. 

What does the law say?

The Family Educational Rights and Privacy Act (FERPA) prohibits the release of Personally Identifiable Information (PII) from student education records without consent or an exception to the requirement for consent as defined by the law (34 CFR Part 99 https://studentprivacy.ed.gov/ferpa).  This means that CSCU employees have legal obligation to avoid releasing information about students that would enable the students to be identified.

According to FERPA, avoiding disclosure of identifiable education records includes more than removing direct identifiers such as name, or date of birth.  The legal definition for PII under FERPA includes “Other information that, alone or in combination, is linked or linkable to a specific student that would allow a reasonable person in the school community, who does not have personal knowledge of the relevant circumstances, to identify the student with reasonable certainty.” (34 CFR §99.3) Therefore, CSCU employees have an obligation to look beyond the first layer of a data product to consider whether student education records could be re-identified and disclosed through linkages to additional information that is being or has been released.

Why is Data Disclosure Avoidance Guidance important?

We live in a data driven world with increasing risks for individuals when personal data is revealed or taken without authorization.  Data about individuals are collected shared and disbursed by public and private organizations regularly and frequently.  So much data is collected, used, and repurposed that individuals are typically unaware of how their personal information is exposed.  Digital profiles of individuals grow silently through the accumulation of additional data points over time, and they are used by actors with both beneficial and harmful intent.  CSCU employees have a legal responsibility to comply with FERPA, and also an ethical responsibility to minimize the additional release of individual data about our students to support student’s privacy rights.  We can do this in part by following this Data Disclosure Avoidance Guidance which provides guardrails for how student education record data should be included in data products produced by our system.  The overall goal is to ensure that all data products released to the public are de-identified to the fullest extent possible.

How do I know if data has been completely de-identified?

De-identification under FERPA is a high standard.  A data set built from student education records is completely de-identified if it doesn’t contain any PII.  According to FEPRA, PII for education records includes “direct identifiers, such as a student’s name or identification number, indirect identifiers, such as a student’s date of birth, or other information which can be used to distinguish or trace an individual’s identity either directly or indirectly through linkages with other information.” (34.CFR §99.3)” The final phrase that PII includes “other information” which can “distinguish or trace” an identity “directly or indirectly” through “linkages with other information, makes it necessary to think beyond the specific data elements in the data product to how those data elements might be used in combination with additional data. https://studentprivacy.ed.gov/content/personally-identifiable-information-education-records). 

The US Department of Education Privacy Technical Assistance Center (PTAC) has supplied the following guidance for determining if data are de-identified.

The challenge is that FERPA was passed into law in 1974, and technical capabilities and interest in linking and leveraging data have evolved dramatically since this time.  With the digital tools and artificial intelligence readily available now it is possible for individuals to be traced and identified when data products are combined to reveal new identifiable characteristics.  Moreover, because data about people are highly valuable, there is enormous interest to capture and combine data of all types.  Organizations around the world regularly scrape publicly available data from the internet to use to make these connections which they either use directly or sell to others.

Since each data product is unique, achieving de-identification requires manual and purposeful review to ensure each data product is de-identified before it is released to non-school officials.

Key Concepts

Cell Suppression

A Cell in a data table refers to an individual piece of information or grouping within the table.  Cells typically hold a count or statistic relevant to the purpose of the table.  For example, a data table with information about enrollments might look like this.  Each ‘box’ is considered a ‘Cell.’  If the value of a file is suppressed, the numeric entry is replaced with a symbol such as an asterisk “*” or a dash “---".

CourseMaleFemaleNon-binaryTotal
Campus 1532678101210
Campus 24214753899
Campus 3*50*65
Complementary Suppression

If the value of a cell is suppressed because the count is below the minimum required for reporting, complementary suppression is employed to ensure that the suppressed value cannot be determined through mathematical manipulation of the data table. With Complementary Suppression the initial cell with a small value is suppressed in addition to the cell with the next highest value in that row or column.  For example, for Campus 3 both of values for Male and Non-binary have been suppressed.  If only one field were suppressed, then the other value could be determined by subtracting from the total. When conducting Complementary Suppression, it is important to consider whether values could be determined from information in both rows and columns for all groups and subgroups when present.   

Personally Identifiable Information (FERPA)

(Title 34, Subtitle A, Part 99, Subpart A, 99.3): The term includes, but is not limited to—

  • The student's name;
  • The name of the student's parent or other family members;
  • The address of the student or student's family;
  • A personal identifier, such as the student's social security number, student number, or biometric record;
  • Other indirect identifiers, such as the student's date of birth, place of birth, and mother's maiden name;
  • Other information that, alone or in combination, is linked or linkable to a specific student that would allow a reasonable person in the school community, who does not have personal knowledge of the relevant circumstances, to identify the student with reasonable certainty; or
  • Information requested by a person who the educational agency or institution reasonably believes knows the identity of the student to whom the education record relates.
Public Release

A Public Release is the release of data to individuals who are not School Officials according to the institution’s FERPA Notice and Directory Information Policy. 

FERPA requires student consent or an allowable exception to consent under the law before student education records with PII can be accessed by individuals who are not School Officials.  Therefore, anyone who is not a School Official, or who is not conducting work related to their professional responsibilities as a School Official is considered a public entity in this context.  Any release to a public entity is a “Public Release.”  Even though CSCU is a system of institutions, each institution (e.g., CT State Community College, Charter Oak, Western CSU) is individually accredited and has separate legal status.  As separate legal entities, individuals employed by one institution are School Officials for that specific institution, not the others; therefore, School Officials at Western may have access to student records for students enrolled at Western, assuming the access is related to their professional responsibilities as a School Official.  These School Officials at Western do not have authorization to access data about students enrolled at CT State without the student’s consent or a FERPA allowable exception to the requirement to obtain consent.  Since the CSCU System Office is the legal governing body over all CSCU institutions, CSCU System Office employees are School Officials of all CSCU institutions and have authorization to access student education records across the system assuming such access is necessary for fulfilling their professional responsibilities. 

Best Practices for Any Data Product

The Fair Information Practice Principles (FIPPs) are an important and trusted framework for guiding the collection and use of data for any data product developed for any audience.  The FIPPs originated in the US Privacy Act of 1974 and are embedded in various forms in major privacy laws around the globe including the European General Data Protection Regulation.  The Federal Privacy Council has a list for federal agencies to follow, https://www.fpc.gov/vision-and-purpose/ .  Another list from the United States Department of Homeland Security is particularly clear and includes the following.  Note, the acronym “DHS” has been replaced with “The agency” for relevance to this guidance. https://www.dhs.gov/publication/privacy-policy-guidance-memorandum-2008-01-fair-information-practice-principles 

Transparency

[The agency] should be transparent and provide notice to the individual regarding its collection, use, dissemination, and maintenance of personally identifiable information (PII).

Individual Participation

 [The agency] should involve the individual in the process of using PII and, to the extent practicable, seek individual consent for the collection, use, dissemination, and maintenance of PII. [The agency] should also provide mechanisms for appropriate access, correction, and redress regarding [The agency’s] use of PII.

Purpose Specification

[The agency] should specifically articulate the authority that permits the collection of PII and specifically articulate the purpose or purposes for which the PII is intended to be used.

Data Minimization

[The agency] should only collect PII that is directly relevant and necessary to accomplish the specified purpose(s) and only retain PII for as long as is necessary to fulfill the specified purpose(s).

Use Limitation

[The agency] should use PII solely for the purpose(s) specified in the notice. Sharing PII outside the Department should be for a purpose compatible with the purpose for which the PII was collected.

Data Quality and Integrity

[The agency] should, to the extent practicable, ensure that PII is accurate, relevant, timely, and complete.
Security: [The agency] should protect PII (in all media) through appropriate security safeguards against risks such as loss, unauthorized access or use, destruction, modification, or unintended or inappropriate disclosure.

Accountability and Auditing

[The agency] should be accountable for complying with these principles, providing training to all employees and contractors who use PII, and auditing the actual use of PII to demonstrate compliance with these principles and all applicable privacy protection requirements.

 

FIPPs which are particularly relevant to the development of data products are “Data Minimization” and “Use Limitation.”  Considerations CSCU employees should take include the following:

  • Understand the purpose of the report
    What does the message of the report say about the groups of students who are the subject of the report?
  • Attend to the sample size
    Are the conclusions about students based on sufficiently large sample sizes to relevant, meaningful, or statistically significant?  Does knowledge of the sample, reveal sensitive personal information?
  • Eliminate small cell sizes
    When producing a data table, are cells with counts less than 10 suppressed? Has the possibility of reverse engineering suppressed values been eliminated by conducting secondary cell suppression and/or using percentages without totals. 
  • Consider using percentages instead of counts and totals
    Omitting totals and substituting percentages for counts minimizes the ability to reverse engineer suppressed counts.
  • Consider using ranges instead of actual counts
    When a report includes information about zero or all individuals (100%) in a group that provides data about everyone in that group.  For example, if a report states that all Hispanic students in Bio 101 received a C- or lower for a semester, then there is no ambiguity about the performance of any student in the class.  Ranges can be used to build in fuzziness. (e.g., zero to 5, or 95-100%)
  • Consider grouping
    Data can be combined across years or categories to enable visibility of patterns for small groups while minimizing identifiability.

DSIR Requirements for Data Disclosure Avoidance

The requirements in this Data Disclosure Guidance differ depending on the use case.  Even after following these guidelines, each report or document with student data needs to be reviewed separately and with consideration for the context in which the student information is reported.  

For data products that are developed by entities outside of CSCU using CSCU student education records, CSCU has the responsibility and authority to review these data products to ensure that all student data is de-identified before they are distributed for Public Release.  No other organization shall determine whether a data product built from CSCU student education records has been de-identified on behalf of CSCU unless that responsibility has been formally delegated in a written agreement.

 

Federally mandated IPEDS reporting

Reporting to the Integrated Postsecondary Education Data System (IPEDS) is mandated for higher education public institutions that receive Title IV funding.  Therefore, reporting to fulfill this requirement is allowed even when counts for specific metrics include individual or very small counts of students.  Most student education data reported to IPEDS are aggregated in some manner; however, it is frequent that institutions are required to report values that pertain to individual students.  For example, an institution might need to report that there was one American Indian or Alaska Native at the institution and additional metrics reveal that this singular individual was a “male,” “degree seeking,” “transfer in,” and “full-time” student.  As another example, it could be possible for individuals in the CSCU community to identify an individual “Asian” student who completed a “Masters” degree for a given year.  Even though required and allowed by law, it is important for CSCU employees to remember that IPEDS data are public and that data points which pertain to individual students could be combined with other data to reveal student identities and reduce privacy for those individuals.

Data products generated by CSCU DSIR for public release

Data that has already been provided to IPEDS may also be disclosed through other reports.  General enrollment and completion information may be reported for any count of individuals.  This includes enrollment information for specific populations (e.g., race/ethnicity, gender, degree, department, enrollment type) as long as the report does not contain multiple characteristics that in combination would lead to identification by a non-school official.

Unless an exception is made by the CSCU Associate Vice President of DSIR, data products containing non- “Directory Information” according to CSCU policy that are provided to any individual or organization outside of system office School Officials shall conform to the following restrictions individually or in combination to minimize the possibility of individual disclosure:

  1. Data at the individual level will not be publicly released.
  2. Suppress all cell counts with values ≤ 5 and consider grouping small categories to increase reportability.
  3. Suppress rates or proportions that are derived from suppressed counts.
  4. Employ complementary cell suppression where necessary.
  5. Use whole numbers when reporting percentages to add ambiguity unless the decimal represents a greater number than 10.
  6. Employ bottom coding taking into consideration the total population size. For example, if the denominator for a statistic is <100, cells in the bottom 10 percent should be recoded as ≤ 10%)
  7. Employ top coding taking into consideration the total population size. For example, if the denominator for a statistic is > 500, cells in the top 5 percent could be recoded as ≥ 95%
  8. Cells with counts of zero, 0% or 100% are be replaced with ranges, see recommendation for bottom and top coding.

For detailed examples of top and bottom coding see the USED, SLDS Technical Brief, “Statistical Methods for Protecting Personally Identifiable Information in Aggregate Reporting

Data products built with CSCU data combined with data sets from others

There are times when CSCU data is joined with data from other sources (e.g., National Student Clearinghouse, College Board, CT Department of Labor, or CT State Department of Education).  When CSCU data is combined with data from other entities, the volume of information and the characteristics about individuals available is increased.  The additional volume of detail dramatically increases the risk of identification; therefore, guidelines for protecting student data are more stringent when data sets are combined.  When CSCU’s data are combined with another agency’s data, the requirements which are most restrictive will be employed. 

  1. Data at the individual level will not be publicly released.
  2. Cells and data points with a count < 10 will be suppressed.
  3. Rates or proportions derived from suppressed counts will also be suppressed.
  4. Complementary cell suppression will be employed so that suppressed values cannot be determined through math or statistics.
  5. Percentages will be presented as whole numbers to add ambiguity.
  6. Statistics must have a denominator ≥
  7. Employ bottom coding taking into consideration the total population size. For example, if the denominator for a statistic is <100, cells in the bottom 10 percent should be recoded as ≤ 10%)
  8. Employ top coding taking into consideration the total population size. For example, if the denominator for a statistic is > 500, cells in the top 5 percent could be recoded as ≥ 95%
  9. Cells with counts of zero, 0% or 100% must be replaced with ranges, see recommendation for bottom and top coding.
  10. Tests of significance that aid in drawing conclusions or projections about student behavior will be based on appropriate sample sizes.
  11. Every data product will be reviewed manually by the developing agency. Automated tools alone will not be relied upon to determine de-identification.
  12. If another organization is creating a data product that was built in part from CSCU student education records, the other organization is required to give CSCU the opportunity, and a minimum of 10 working days, to review the data products to ensure they are completely de-identified before Public Release. No other organization can make this determination for CSCU unless it has been authorized to do so through written agreement.
  13. Case-by-case exceptions to suppression rules may be made with approval from the Associate Vice President of DSIR.

Approval & Revision History

Approved by

Nancy Becerra-Cordoba
Assistant Vice President
Decision Support and Institutional Research    

Date
6/13/24

 

Revision History
Revision DateVersionDescription / ReasonRevised ByApproved ByApproval Date
2025-11-11V2Change “Policy” to “Guidance” and small grammar correctionsJan KiehneJan Kiehne11-11-2025