Algorithmic Tools Compliance Report

"This is an annual report on algorithmic tools used by City agencies, collected under Local Law 35 of 2022 (LL 35). It includes descriptions of the tool's use and purpose, datasets used, and vendor involvement. The report is also published as a PDF on the OTI website: https://www.nyc.gov/content/oti/pages/reports.

The full text of LL 35 is available online: https://legistar.council.nyc.gov/LegislationDetail.aspx?ID=4265421&GUID=FBA29B34-9266-4B52-B438-A772D81B1CB5

An "algorithmic tool" is defined by the law as: "Any technology or computerized process that is derived from machine learning, artificial intelligence, predictive analytics, or other similar methods of data analysis, that is used to make or assist in making decisions about and implementing policies that materially impact the rights, liberties, benefits, safety or interests of the public, including their access to available city services and resources for which they may be eligible. Such term includes, but is not limited to tools that analyze datasets to generate risk scores, make predictions about behavior, or develop classifications or categories that determine what resources are allocated to particular groups or individuals, but does not include tools used for basic computerized processes, such as calculators, spellcheck tools, autocorrect functions, spreadsheets, electronic communications, or any tool that relates only to internal management affairs such as ordering office supplies or processing payments, and does not materially affect the rights, liberties, benefits, safety or interests of the public."

City Government Office of Technology and Innovation (OTI) Dataset jaw4-yuem 27 fields
DOWNLOAD CSV
Dataset fields
Showing 50 real records
Department of Health and Mental Hygiene
Year: 2024 • Agency: Department of Health and Mental Hygiene • Department: Disease Control - Public Health Laboratory
Year
2024
Agency
Department of Health and Mental Hygiene
Department
Disease Control - Public Health Laboratory
Tool Name
kSNP4
Date First Use
2022/03
Updated
Yes
Purpose Type
Data management
Computation Type
Clustering
Autonomy
NA
Frequency
NA
Population Type
Individuals; Biological sample
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
kSNP4 can use multiple algorithms (e.g., maximum-likelihood, parsimony, neighbor-joining) to infer phylogenetic trees from genomes.
Purpose Desc
Produced phylogenetic trees are used to help rule in or out outbreaks of bacteria.
Updated Desc
Previously used kSNP3, updated this year to kSNP4.
Identifying Info
false
Data Training
N/A
Data Input
Fasta
Data Output
ML tree in NEWICK format, log & configuration files, Fasta file.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
NA
Department of Health and Mental Hygiene
Year: 2024 • Agency: Department of Health and Mental Hygiene • Department: Disease Control - Public Health Laboratory
Year
2024
Agency
Department of Health and Mental Hygiene
Department
Disease Control - Public Health Laboratory
Tool Name
Minimap2
Date First Use
2020/05
Updated
No
Purpose Type
Data management
Computation Type
Matching
Autonomy
NA
Frequency
NA
Population Type
Individuals; Biological sample
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Aligns sequencing data to a reference sequence.
Purpose Desc
Minimap2 uses optimal chaining scores to align sequencing data to reference genomes. This tool is faster and more optimal for long read sequences, such as Oxford Nanopore Technologies data. This tool is used to predict the order in which the fragments generated by sequencers are pieced together to form a complete genomic sequence data. This tool is used for COVID-19 and monkeypox virus sequencing analyses.
Updated Desc
NA
Identifying Info
false
Data Training
N/A
Data Input
Sequence reads (fastq) for single or paired-end runs (sequence reads can be considered strings).
Data Output
Aligned reads in SAM format.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
NA
Department of Health and Mental Hygiene
Year: 2024 • Agency: Department of Health and Mental Hygiene • Department: Disease Control - Public Health Laboratory
Year
2024
Agency
Department of Health and Mental Hygiene
Department
Disease Control - Public Health Laboratory
Tool Name
Multiple Alignment using Fast Fourier Transform (MAFFT)
Date First Use
2021/01
Updated
No
Purpose Type
Data management
Computation Type
Matching
Autonomy
NA
Frequency
NA
Population Type
Individuals; Biological sample
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Aligns multiple sequencing data.
Purpose Desc
Multiple Alignment using Fast Fourier Transform (MAFFT) includes several algorithmic methods, including guided tree, scoring matrices, and sequence alignment algorithms to realign multiple genomic sequencing data. The tool aligns sequences, to help identify differences. This is used in all sequencing analysis prior to building a phylogenetic tree or distance tree.
Updated Desc
NA
Identifying Info
false
Data Training
N/A
Data Input
Sequences can be in GCG, FASTA, EMBL (nucleotide only), GenBank, PIR, NBRF, PHYLIP or UniProtKB/Swiss-Prot (protein only) format.
Data Output
FASTA or Clustalw.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
NA
Department of Health and Mental Hygiene
Year: 2024 • Agency: Department of Health and Mental Hygiene • Department: Disease Control - Public Health Laboratory
Year
2024
Agency
Department of Health and Mental Hygiene
Department
Disease Control - Public Health Laboratory
Tool Name
Pangolin
Date First Use
2021/07
Updated
Yes
Purpose Type
Data management
Computation Type
Clustering
Autonomy
NA
Frequency
NA
Population Type
Individuals; Biological sample
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Assigns lineage names to SARS-CoV-2.
Purpose Desc
Pangolin uses a combination of several methods, including random forest tree, classification methods, and maximum parsimony to assign lineage names to SARS-CoV-2 genomic sequences to bin sequences that are more likely to be similar. This is a tool that designates a name based on a nomenclature for COVID-19 sequence data.
Updated Desc
Pangolin is updated as new sequences are published, and as the virus evolves. Currently, it is version 4.3.
Identifying Info
false
Data Training
Trained on a dataset of genomes that have been designated to Pango lineages using whole genome information.
Data Input
Fasta files.
Data Output
.csv file with taxon name and lineage assigned.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
NA
Department of Health and Mental Hygiene
Year: 2024 • Agency: Department of Health and Mental Hygiene • Department: Disease Control - Public Health Laboratory
Year
2024
Agency
Department of Health and Mental Hygiene
Department
Disease Control - Public Health Laboratory
Tool Name
PHYLOViZ
Date First Use
2017/10
Updated
No
Purpose Type
Data management
Computation Type
Clustering
Autonomy
NA
Frequency
NA
Population Type
Individuals; Biological sample
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
For representing the possible evolutionary relationships between strains, PHYLOViZ uses the goeBURST algorithm, a refinement of eBURST algorithm by Feil et al., and its expansion to generate a complete minimum spanning tree.
Purpose Desc
Used to generate the minimum spanning tree relationships.
Updated Desc
NA
Identifying Info
false
Data Training
N/A
Data Input
txt, NEWICK, FASTA.
Data Output
Minimum spanning tree.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
NA
Department of Health and Mental Hygiene
Year: 2024 • Agency: Department of Health and Mental Hygiene • Department: Disease Control - Public Health Laboratory
Year
2024
Agency
Department of Health and Mental Hygiene
Department
Disease Control - Public Health Laboratory
Tool Name
PulseNet 2.0
Date First Use
2017/09
Updated
Yes
Purpose Type
Data management
Computation Type
Clustering
Autonomy
NA
Frequency
NA
Population Type
Individuals; Biological sample
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
A suite of tools used to align and analyze bacterial genomes.
Purpose Desc
PulseNet 2.0 is used to:
1. Assemble the bacterial genome (since the sequencing process involves fragmenting the bacterial DNA and then amplifying it into millions of pieces);
2. Identify the genus, species, and serotype of the bacterial isolate;
3. Perform quality control checks to ensure the sequence meets certain quality standards;
4. Perform core and whole genome multi-locus sequence typing (a technique used to type bacteria based on their genetic code);
5. Perform cluster analysis for cases related to one another based upon case definitions recommended by the Centers for Disease Control and Prevention (CDC).
This information is then communicated to partners including foodborne epidemiologists at the Bureau of Communicable Disease, who investigate all reported cases of foodborne disease, with those investigations potentially resulting in restaurant inspections, closures, and food recalls.
Updated Desc
This tool was put on the cloud and accessed through CDC Secure Access Management Services site.
Identifying Info
false
Data Training
N/A
Data Input
Fastq.
Data Output
txt, Excel.
Vendor Name
Centers for Disease Control and Prevention
Vendor Type
NA
Vendor Desc
Developed and maintains the tool.
Data 2022
NA
Vendor
NA
Analysis Type
NA
Department of Health and Mental Hygiene
Year: 2024 • Agency: Department of Health and Mental Hygiene • Department: Disease Control - Public Health Laboratory
Year
2024
Agency
Department of Health and Mental Hygiene
Department
Disease Control - Public Health Laboratory
Tool Name
Spades
Date First Use
2017/10
Updated
No
Purpose Type
Data management
Computation Type
Matching
Autonomy
NA
Frequency
NA
Population Type
Individuals; Biological sample
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Spades uses several algorithms to simplify genomic read data into de Brujin graphs and finds overlaps to assemble genomes.
Purpose Desc
Spades is an intermediate step in the workflows of bacterial analyses.
Updated Desc
NA
Identifying Info
false
Data Training
N/A
Data Input
Fastq.
Data Output
Fastas and other files for corrected reads; scaffolds, contigs, paths in GFA format; fastg assembly graph.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
NA
Department of Health and Mental Hygiene
Year: 2024 • Agency: Department of Health and Mental Hygiene • Department: Disease Control - Public Health Laboratory
Year
2024
Agency
Department of Health and Mental Hygiene
Department
Disease Control - Public Health Laboratory
Tool Name
Vsearch
Date First Use
2022/06
Updated
No
Purpose Type
Data management
Computation Type
Matching
Autonomy
NA
Frequency
NA
Population Type
Individuals; Biological sample
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Vsearch uses the Needleman-Wunsch algorithm to merge read pairs and align and dereplicate sequences to detect chimeric genomic sequences.
Purpose Desc
Vsearch is an intermediate step in the workflow to analyze COVID-19 variants in wastewater.
Updated Desc
NA
Identifying Info
false
Data Training
N/A
Data Input
Sequence reads (fastq, Fasta) for single or paired-end runs (sequence reads can be considered strings).
Data Output
FASTA, FASTQ, tables, alignments, SAM.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
NA
Department of Investigation
Year: 2024 • Agency: Department of Investigation • Department: NA
Year
2024
Agency
Department of Investigation
Department
NA
Tool Name
Facial Recognition Technology
Date First Use
2019/03
Updated
Yes
Purpose Type
Not specified
Computation Type
Matching
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The tool analyzes an uploaded image or video and searches and compares it with lawfully possessed images to generate a pool of possible matches. If possible matches are identified, a trained DOI examiner visually analyzes and evaluates potential matches to assess reliability of a match consistent with agency policy and applicable laws. A match serves as an investigative lead for additional investigative steps and does not constitute a positive identification.
Purpose Desc
Facial recognition is a digital technology that DOI uses to analyze uploaded images or videos of people and objects obtained during an investigation by comparison with lawfully possessed images. Facial recognition generates possible matches of an object or individual from this analysis and comparison. The purpose of the tool is to assist DOI investigations of matters within its jurisdiction including fraud and other criminal activity.
Updated Desc
General system updates by vendor to improve quality of search results and to fix bugs.
Identifying Info
true
Data Training
Vendor uses publicly available open source media data.
Data Input
Images.
Data Output
Images.
Vendor Name
Not disclosable
Vendor Type
NA
Vendor Desc
Out-of-the-box products. The vendors provide ongoing technical assistance. Confidentiality agreements are in place with the vendors.
Data 2022
NA
Vendor
NA
Analysis Type
NA
Department of Social Services
Year: 2024 • Agency: Department of Social Services • Department: NA
Year
2024
Agency
Department of Social Services
Department
NA
Tool Name
Homebase Risk Assessment Questionnaire (RAQ)
Date First Use
2012/06
Updated
No
Purpose Type
Resource allocation
Computation Type
Scoring
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
Households seeking Homebase assistance
Population Type Other
NA
Website
NA
Tool Desc
Homebase applicants answer screening questions about their current housing situation, history of disruptive experiences, shelter history, and other domains. Each of the answers is assigned a number of points, and applicants that reach a certain point threshold are eligible for deeper Homebase services, such as financial assistance and case management. Workers are able to override a limited number of model decisions with permission of a supervisor.
Purpose Desc
The Homebase program was created to prevent households from entering the DHS shelter system. Since NYC has a range of antipoverty programs and the number of households entering shelter is small compared to the pool of New Yorkers who have an eviction filing each year, the Agency had to ensure that the households who most needed additional homelessness prevention services were being enrolled in Homebase programs. Research showed that staff were not accurately able to predict who would or would not enter the DHS shelter system and that using a risk assessment would provide a better way to match resources to the families who would benefit the most.
Updated Desc
NA
Identifying Info
true
Data Training
The RAQ was developed based on analysis of data on Homebase enrollees from 2004 to 2008, conducted in conjunction with a team of academic researchers, to determine predictive factors for those entering shelter. It was updated in 2023 based on analysis led by DSS researchers of 2013-2016 Homebase data.
Data Input
Factors include, among others: personal characteristics such as age and pregnancy; educational attainment and employment status; housing issues such as eviction, discord, a move in the past year; past and recent experience of homelessness.
Data Output
The tool produces a score that is used to assess eligibility for full versus brief Homebase services.
Vendor Name
Multiple researchers
Vendor Type
NA
Vendor Desc
DHS contracted with researchers to evaluate years of Homebase administrative data to develop a risk assessment. The DSS research team then led an updated analysis that led to tool revisions. The published research papers are listed below:

https://ajph.aphapublications.org/doi/10.2105/AJPH.2013.301468

https://www.journals.uchicago.edu/doi/abs/10.1086/686466?mobileUi=0&journalCode=ssr

https://www.tandfonline.com/doi/abs/10.1080/10511482.2022.2077801
Data 2022
NA
Vendor
NA
Analysis Type
NA
Department of Social Services
Year: 2024 • Agency: Department of Social Services • Department: Public Engagement Unit
Year
2024
Agency
Department of Social Services
Department
Public Engagement Unit
Tool Name
NYC Enterprise Data Solutions Service (EDS)
Date First Use
2024/03
Updated
Created in CY2024
Purpose Type
Data management
Computation Type
Matching
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The NYC Enterprise Data Solutions Service (EDS) automatically matches and links person records from multiple source systems, regardless of whether they share a common identifier such as a social security number. EDS supports cross-agency data integration and interoperability, particularly across health and human service agencies. EDS works by processing business rules that standardize name and geocode address data before applying a series of deterministic and probabilistic matching rules that generate detailed or summarized reports of record linkages.
Purpose Desc
In 2024, the Mayor’s Public Engagement Unit (PEU) moved the operations of its Rent Freeze team (focused on enrollment in Senior Citizen Rent Increase Exemption, Disability Rent Increase Exemption, Senior Citizen Homeowners’ Exemption, and Disabled Homeowners’ Exemption programs) to a new case/client management system within our custom Salesforce instance. Prior to migrating data from the prior system (based in our EveryAction tool), PEU needed to identify duplicate clients within EveryAction and between EveryAction and Salesforce. The primary reason for duplication is that we use Salesforce to support a number of other teams that might have interacted with Rent Freeze clients as well. We used the EDS tool developed by our partners at NYC Opportunity and maintained by OTI to identify client records that were highly likely to be the same client. PEU developed its own process to combine outreach and case history information for the records identified as related to the same client to create a single record in PEU’s Salesforce system across programs. Using the EDS tool helped us provide better service to our clients by providing our specialists with richer information about the client’s interaction with the Rent Freeze, Tenant Helpline, Tenant Support Unit, and GetCovered teams.
Updated Desc
NA
Identifying Info
true
Data Training
N/A
Data Input
Client contact information, including name, phone numbers, email addresses, home addresses.
Data Output
The same client information but with common identifier numbers.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
NA
Department of Social Services
Year: 2024 • Agency: Department of Social Services • Department: Public Engagement Unit
Year
2024
Agency
Department of Social Services
Department
Public Engagement Unit
Tool Name
SmartVAN / TargetSmart
Date First Use
2019/11
Updated
Yes
Purpose Type
Data management
Computation Type
Matching
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The Mayor’s Public Engagement Unit (PEU) uses SmartVAN to manage outreach across a range of projects. SmartVAN provides functionality to create lists of potential clients to contact, collect personal information and survey responses from clients, and conduct outreach via phone banks and canvassing. SmartVAN also contains a frequently updated commercial dataset, provided by TargetSmart, of New York City residents and their demographic, contact, and other information. PEU uses this preloaded data to create outreach lists when data on existing clients or from partner agencies is unavailable or insufficient to meet the scope of the outreach project.
Purpose Desc
In 2023, PEU had the TargetSmart data within SmartVAN on a number of projects. PEU frequently uses the data to create lists of residents who live within certain zip codes that PEU wants to target for outreach. For example, PEU created lists using TargetSmart data to conduct door-knocking and phone banking outreach to New Yorkers identified as potentially eligible for the DOF Rent Freeze program based on TargetSmart data. In cases like these, TargetSmart’s determination of who lives in which zip codes as well as estimated income affects whether New Yorkers receive PEU outreach. Additionally, the algorithm that TargetSmart uses to match phone numbers to individuals impacts the type of outreach that New Yorkers receive.
Updated Desc
We get regular updated lists of New Yorkers from the vendor directly into the EveryAction platform we use for outreach.
Identifying Info
true
Data Training
Training data is part of vendor’s proprietary processes.
Data Input
Input data is part of vendor’s proprietary processes.
Data Output
The algorithmically derived data that PEU accesses is the output of proprietary algorithmic processes developed and operated by TargetSmart. These algorithmic processes include matching multiple input datasets to determine residency, contact information, and demographics on New York City residents. SmartVAN also includes a number of algorithmically determined likelihood scores, including scores for the likelihood that a household contains children under 18, etc.
Vendor Name
EveryAction and TargetSmart
Vendor Type
NA
Vendor Desc
EveryAction and TargetSmart jointly provide the SmartVAN product. EveryAction is the software provider. TargetSmart is the data provider. TargetSmart is the entity who applies algorithmic techniques. EveryAction provides access to this data through their platform.
Data 2022
NA
Vendor
NA
Analysis Type
NA
Department of Transportation
Year: 2024 • Agency: Department of Transportation • Department: Traffic Operations
Year
2024
Agency
Department of Transportation
Department
Traffic Operations
Tool Name
Midtown in Motion / Adaptive Control Decision Support System
Date First Use
2011/07
Updated
Yes
Purpose Type
Performance evaluation
Computation Type
Classification
Autonomy
NA
Frequency
NA
Population Type
Geographic space; Individuals
Population Type Individual
Travelers in the Manhattan midtown core
Population Type Other
NA
Website
NA
Tool Desc
Midtown in Motion (MIM) is a program that measures traffic congestion in the midtown core of Manhattan (from 1st to 9th Avenues and from 57th to 34th Streets, inclusive), using sensor data captured within and beyond the zone. This congestion classification (light, moderate, moderate-heavy, heavy) is used by the Adaptive Control Decision Support System (ACDSS) to choose the optimal signal timing along the avenues to reduce congestion and improve traffic flow for all modes of transportation.
Purpose Desc
The purpose of the tool is to improve congestion in the midtown core of Manhattan through responsive traffic signal management.
Updated Desc
Ongoing maintenance.
Identifying Info
false
Data Training
N/A
Data Input
Motor vehicle travel times captured by Radio Frequency Identification (RFID) sensors.
Data Output
Recommended traffic signal plan.
Vendor Name
KLD Engineering , P.C.
Vendor Type
NA
Vendor Desc
KLD Engineering, P.C. is the developer of ACDSS and currently handles the maintenance contract.
Data 2022
NA
Vendor
NA
Analysis Type
NA
Fire Department
Year: 2024 • Agency: Fire Department • Department: Management Analysis and Planning
Year
2024
Agency
Fire Department
Department
Management Analysis and Planning
Tool Name
Emergency Medical Services (EMS) Hospital Suggestion Algorithm
Date First Use
2021/03
Updated
No
Purpose Type
Resource allocation
Computation Type
Ranking
Autonomy
NA
Frequency
NA
Population Type
Individuals; Geographic space; Group, organization, or business
Population Type Individual
Patients
Population Type Other
NA
Website
NA
Tool Desc
The Emergency Medical Services (EMS) Hospital Suggestion Algorithm is used to determine the closest, appropriate hospital to the incident location based on the medical needs of a patient requiring transport.
Purpose Desc
The algorithm computes a list of hospitals in order of closest to furthest in time for each medical condition category as currently established. (For example, there is a list of hospitals computed in order of closest in time for all hospitals that accept General Emergency Department patients and for all hospitals that accept special conditions, such as burns). Depending on the medical needs category of the patient, the algorithm produces a pre-determined list of hospitals which is based on the location of the patient and then made available to the crew as a list of “closest, most appropriate hospitals.”
Updated Desc
NA
Identifying Info
false
Data Training
The EMS Hospital Suggestion algorithm relies on automatic vehicle location data from ambulances transporting patients to hospitals between 2018 and 2019 to calibrate a network analysis model that derives incident to hospital transport times. The order of suggested hospitals are then compared with five years of historical EMS hospital transport data from before the COVID-19 pandemic (2015-2019) to validate and correct the network model.
Data Input
The inputs for the algorithm include the location and medical call type of the patient.
Data Output
The algorithm outputs the closest, most appropriate hospitals.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
NA
Fire Department
Year: 2024 • Agency: Fire Department • Department: Management Analysis and Planning
Year
2024
Agency
Fire Department
Department
Management Analysis and Planning
Tool Name
Emergency Medical Services (EMS) Unit Suggestion Algorithm
Date First Use
2007/03
Updated
No
Purpose Type
Resource allocation
Computation Type
Ranking
Autonomy
NA
Frequency
NA
Population Type
Geographic space; Individuals; Group, organization, or business
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The Emergency Medical Services (EMS) Unit Suggestion Algorithm is used to determine which order of geographic regions (known as atoms) to search in order for the EMSCAD system to select an appropriate EMS unit for dispatch to an incident.
Purpose Desc
The algorithm computes a list of geographic regions (known as atoms) in order of closest to furthest in travel time for each atom in the city. This list of ordered atoms is the output of an algorithm that relies on a calibrated network model to derive travel time estimates. The output is an excel file which is converted into an EMSCAD-compatible file and loaded into the system for real-time unit selection capabilities. The file is generated and implemented as a 24/7 source file, meaning, the recommended search order is not currently varying by time of day. The department is intending to implement time-of day search orders in the near future.
Updated Desc
NA
Identifying Info
false
Data Training
The EMS Unit Suggestion algorithm relies on historical EMSCAD trip time data which is used to calibrate a network analysis model which derives atom-to-atom transport times.
Data Input
The input for the algorithm is a geographic location.
Data Output
The algorithm outputs a recommended EMS unit for dispatch.
Vendor Name
Deccan International
Vendor Type
NA
Vendor Desc
This algorithm and the resulting output file that is used in our EMSCAD system to suggest atom order for unit search is currently provided by a vendor, Deccan International.
Data 2022
NA
Vendor
NA
Analysis Type
NA
Fire Department
Year: 2024 • Agency: Fire Department • Department: Management Analysis and Planning
Year
2024
Agency
Fire Department
Department
Management Analysis and Planning
Tool Name
RBIS (Risk Based Inspection Program); ALARM (A Learning Approach to Risk Modeling)
Date First Use
2019/11
Updated
Yes
Purpose Type
Risk management
Computation Type
Scoring
Autonomy
NA
Frequency
NA
Population Type
Property; Individuals; Group, organization, or business
Population Type Individual
Civilian fire injuries/fatalities
Population Type Other
NA
Website
NA
Tool Desc
A Learning Approach to Risk Modeling (ALARM) creates risk scores for each building in the city. These scores are used to schedule our Fire Operations building inspections within the inspectable population of buildings in the city (~330,000 Building Identification Numbers), as a part of the Risk-Based Inspection Program (RBIS).
Purpose Desc
ALARM is a combined approach using machine learning and risk ratios to assess the risk of a building for structural fire ignition (probability) and civilian fire injury/death (impact). The machine learning algorithm takes incident data, housing characteristics, and NYC311 data and creates a probability of structural fire ignition. This is combined with a civilian injury or death risk ratio for the building which is based on building characteristics, incident data and nearby felony crimes to create a risk score (range is one to nine), with one being highest risk and nine being lowest risk. Buildings are prioritized within each of the nine risk scores according to the residential population in each building.
Updated Desc
The fire ignition and injury/death models are recalculated monthly with fresh training data to create updated variable weights.
Identifying Info
false
Data Training
Each month the team uses a five-year incident dataset and reserves 99 percent of the data to train the ignition model and 80 percent of the data to train the impact model.
Data Input
The ALARM risk score utilizes data from our Fire and Emergency Medical Services dispatch systems, building characteristic data, NYC311 calls, felony crimes, census data and civilian injury data.
Data Output
The tool outputs a risk score from one (highest risk) to nine (lowest risk).
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
NA
Mayor's Office
Year: 2024 • Agency: Mayor's Office • Department: Mayor's Office of Media and Entertainment
Year
2024
Agency
Mayor's Office
Department
Mayor's Office of Media and Entertainment
Tool Name
Adobe Photoshop
Date First Use
2024/01
Updated
No
Purpose Type
Information presentation
Computation Type
Data generation
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
Viewers who watch the city’s television channels; some individuals who appear in Mayor’s Office of Media and Entertainment’s television content
Population Type Other
NA
Website
NA
Tool Desc
NYC Media uses Adobe Photoshop to edit images. Adobe Photoshop uses generative AI to allow users to edit images without manual work.
Purpose Desc
NYC Media uses Adobe Photoshop to make slight edits to some images that appear in some content that is produced in-house and broadcast on the city’s television edits. For example, to comply with Federal Communications Commission regulations for non-commercial educational stations, we may use an AI tool to blur a company’s logo on a t-shirt. As another example, we may use the AI tool to add visual interest, for example, to add legs in a picture that is cropped at the waist.
Updated Desc
NA
Identifying Info
true
Data Training
According to Adobe’s website, “generative AI Image models were trained on licensed content, such as Adobe Stock, and public domain content where copyright has expired.”
Data Input
Images copyrighted by the City of New York or licensed to the City of New York pursuant to an agreement that authorizes edits and, if involving images of people, content that is covered by a written consent form.
Data Output
Visual content that is broadcast on the City’s television stations.
Vendor Name
Adobe
Vendor Type
NA
Vendor Desc
Adobe regularly updates the Photoshop software.
Data 2022
NA
Vendor
NA
Analysis Type
NA
Mayor's Office
Year: 2024 • Agency: Mayor's Office • Department: Mayor's Office of Media and Entertainment
Year
2024
Agency
Mayor's Office
Department
Mayor's Office of Media and Entertainment
Tool Name
Adobe Premiere Pro
Date First Use
2021/00
Updated
No
Purpose Type
Information presentation
Computation Type
Data generation
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
Individuals who watch content on NYC Media's television channels; individuals who speak in content on NYC Media's television channels
Population Type Other
NA
Website
NA
Tool Desc
NYC Media uses Adobe Premiere Pro to edit video content broadcast on the city’s television stations. Within Adobe Premiere Pro, we use AI-powered tools to help generate closed captions of some video content that is edited in-house prior to broadcast.
Purpose Desc
We use AI-powered tools to help a human editor generate closed captions of some video content that is edited in-house prior to broadcast on the city’s television channels.
Updated Desc
NA
Identifying Info
true
Data Training
According to Adobe’s website, “Speech to Text is powered by a combination of Adobe proprietary technology — including Adobe Sensei machine learning— and third-party technologies.”
Data Input
Spoken words in video programs.
Data Output
Closed captions.
Vendor Name
Adobe
Vendor Type
NA
Vendor Desc
Adobe provides regular software updates.
Data 2022
NA
Vendor
NA
Analysis Type
NA
Mayor's Office
Year: 2024 • Agency: Mayor's Office • Department: Mayor’s Office to End Domestic and Gender Based Violence
Year
2024
Agency
Mayor's Office
Department
Mayor’s Office to End Domestic and Gender Based Violence
Tool Name
AI Transcription on Teams
Date First Use
2024/04
Updated
Created in CY2024
Purpose Type
Information presentation
Computation Type
Data generation
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
Participants in the meeting might provide their name and business affiliations
Population Type Other
NA
Website
NA
Tool Desc
Microsoft Teams has a build in feature that uses AI to create a transcript of the meeting.
Purpose Desc
The tool provided a transcript of a meeting which was then reviewed by Mayor’s Office to End Domestic and Gender Based Violence staff for accuracy. The transcript was then emailed to meeting participants.
Updated Desc
NA
Identifying Info
true
Data Training
Training data is proprietary to Microsoft.
Data Input
The words spoken during a meeting are captured by the tool.
Data Output
The tool provides text of the words spoken during the meeting.
Vendor Name
Microsoft
Vendor Type
NA
Vendor Desc
Microsoft provides this tool as part of Teams.
Data 2022
NA
Vendor
NA
Analysis Type
NA
Mayor's Office
Year: 2024 • Agency: Mayor's Office • Department: Mayor's Office of Media and Entertainment
Year
2024
Agency
Mayor's Office
Department
Mayor's Office of Media and Entertainment
Tool Name
AppTek OmniCaption 300 Closed Captioning Appliance
Date First Use
2022/11
Updated
No
Purpose Type
Information presentation
Computation Type
Data generation
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
People who watch live City Council and mayoral content televised on NYC Gov and other content televised on NYC World; people who appear in the content that is televised
Population Type Other
NA
Website
NA
Tool Desc
The tool uses AI-enabled automatic speech recognition to create closed captions of live television content.
Purpose Desc
The Mayor’s Office of Media and Entertainment is using the AppTek Omni 300 closed captioning appliance to provide closed captioning of live-broadcasted events (e.g., City Council hearings) and content that is cablecast on NYC World.
Updated Desc
NA
Identifying Info
true
Data Training
According to AppTek’s website, the “OmniCaption 300 closed captioning appliance was developed for and trained on broadcast news, sports, weather and other programming.”
Data Input
The input data are words spoken by people during live broadcasts of public hearings, meetings, and events and content on NYC World.
Data Output
Closed captions that reflect the written text of the input data (spoken words).
Vendor Name
AppTek
Vendor Type
NA
Vendor Desc
AppTek provides support for the Omni 300 closed captioning system.
Data 2022
NA
Vendor
NA
Analysis Type
NA
Mayor's Office
Year: 2024 • Agency: Mayor's Office • Department: Mayor’s Office to End Domestic and Gender Based Violence
Year
2024
Agency
Mayor's Office
Department
Mayor’s Office to End Domestic and Gender Based Violence
Tool Name
ChatGPT
Date First Use
2024/03
Updated
No
Purpose Type
Information presentation
Computation Type
Data generation
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
ChatGPT is an advanced AI language model that can understand and generate human-like text based on the input it receives. It can assist with a wide range of tasks, including answering questions, providing recommendations, and engaging in meaningful conversations.
Purpose Desc
The Mayor’s Office to End Domestic and Gender Based Violence (ENDGBV) staff have used ChatGPT to complete the following tasks: assist in gathering information for literature reviews for public facing ENDGBV reports and assist in creating potential job interview questions based on content in the job description.
Updated Desc
NA
Identifying Info
false
Data Training
Training data is proprietary to OpenAI.
Data Input
Text prompts were provided to ChatGPT.
Data Output
ChatGPT provided text responses to the text prompts.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
NA
Mayor's Office
Year: 2024 • Agency: Mayor's Office • Department: Civic Engagement Commission
Year
2024
Agency
Mayor's Office
Department
Civic Engagement Commission
Tool Name
Methodology for Poll Site Language Assistance
Date First Use
2020/11
Updated
Yes
Purpose Type
Resource allocation
Computation Type
Ranking
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Since no dataset is currently available that reliably captures the number of limited English proficient (LEP) registered voters for all program languages, the Civic Engagement Commission (CEC) uses the percentage of LEP citizens of voting age (CVALEP) as a substitute or proxy measure of need. CEC ranks the program-eligible languages in order of magnitude of CVALEP and distributes poll sites to each language based on its ranking (excluding CVALEP persons that speak languages served by the NYC Board of Elections in certain New York City counties). The number of poll sites that will receive services in any given language will depend on each language’s share of the total CVALEP in the population eligible to be served. For example, according to U.S. Census data, approximately 207,926 New Yorkers are CVALEP and speak a language that is served by this program. This proportionality approach allows CEC to balance goals of including diverse language communities as well as fair access to the total number of eligible voters within each language community. The program provides interpreters in program-eligible languages at poll sites based on U.S. Census data showing concentrations of CVALEP individuals who speak these languages and reside around each poll site. For each language, poll sites are chosen in descending order of concentration of CVALEP, until the language’s share is met. This process is repeated for each language, thereby including the poll sites with the highest concentration of CVALEP for each program-eligible language until that language’s share is met, and the total number of poll sites for which resources are allocated is reached. It may be possible, based on analysis of data, to reassign poll sites to languages with greater need; however, each language will receive a minimum of at least one poll site. Models used included the Thiessen polygon method to create a Voronoi diagram to determine CVALEP estimates.
Purpose Desc
This is a methodology for determining how the CEC will provide interpretation services at poll sites for LEP voters. The methodology explains how the CEC will identify the languages and locations in which interpretation services will be offered during the November 2024 election and beyond. These services supplement the interpretation assistance provided by NYC Board of Elections in several languages. Under the Charter, the CEC can only provide interpretation services in a language if it is a designated citywide language or it is spoken by a greater number of LEP New Yorkers than the lowest ranked designated citywide language and at least one poll site has a significant concentration of speakers of such language with LEP. This methodology ensures service for all languages that are eligible under the Charter.
Updated Desc
We are now using 2017-2021 data from the American Community Survey 5-year data. The differences between new and previous data are not statistically significant, therefore the distribution of services will not be affected based on these differences.
Identifying Info
false
Data Training
N/A
Data Input
For citywide estimates, this methodology uses current data from the American Community Survey 2017-2021 5-year estimates. This methodology also uses the American Community Survey Census Tract 2017-2021 5-year Public Use Microdata Samples for poll site level analysis, which tracks resident New Yorkers at the neighborhood level. In addition, the methodology uses data from the Board of Elections on the location of election districts and poll sites.
Data Output
The algorithm estimates the number of citizens of voting age with Limited English Proficiency for each program-eligible language who could report to each polling site.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
NA
Mayor's Office
Year: 2024 • Agency: Mayor's Office • Department: Civic Engagement Commission
Year
2024
Agency
Mayor's Office
Department
Civic Engagement Commission
Tool Name
StratifySelect
Date First Use
2022/11
Updated
No
Purpose Type
Resource allocation
Computation Type
Matching
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
New York City residents (all boroughs)
Population Type Other
NA
Website
NA
Tool Desc
StratifySelect is used to select a group of people from a pool of applicants such that the selected group matches target demographics and is as randomized as possible given the need to match targeted demographics. Traditional stratified sampling can create cases in which some individuals have a near-zero chance of being selected. This method uses a new technique of explicitly computing a maximally fair output distribution and then sampling from that distribution to select the final panel, achieving a fairer distribution of probabilities per applicant while maintaining fidelity to the demographics of the borough. The tool compares applicant demographic data to the data from the American Community Survey to achieve a representative sample. See this paper in Nature for details on the methodology: https://www.nature.com/articles/s41586-021-03788-6
Purpose Desc
The group selected by StratifySelect will be invited to participate in the CEC’s Borough Assemblies. Each group contains a prioritized list including “backup options” in case some of the invitees decline or are unable to participate.
Updated Desc
NA
Identifying Info
true
Data Training
N/A
Data Input
Fully anonymized data from interested people (age, gender, race, and Hispanic identity, borough, zip code) is used as input. We emphasize that no identifying data is stored, shared and/or transmitted during the entire process.
Data Output
The output is a subset of the same data such that the output group is both randomly selected and representative of each borough in these categories, as defined by public census data.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
NA
Mayor's Office
Year: 2024 • Agency: Mayor's Office • Department: Mayor's Office of Media and Entertainment
Year
2024
Agency
Mayor's Office
Department
Mayor's Office of Media and Entertainment
Tool Name
Zoom
Date First Use
2020/00
Updated
No
Purpose Type
Information presentation
Computation Type
Data generation
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
People who participate in rulemaking hearings and webinars; people who read transcripts of those hearings and webinars
Population Type Other
NA
Website
NA
Tool Desc
Zoom is a virtual meeting platform; Zoom has an auto closed-caption function that uses AI.
Purpose Desc
The Mayor’s Office of Media and Entertainment (MOME) uses Zoom for public hearings on rulemaking and for public webinars. We use Zoom’s auto transcript function and captioning function. (Note: We provide American Sign Language and human-typed Communication Access Realtime Translation services as a reasonable accommodation upon request.) If we publish a transcript after the Zoom meeting, a human reviews and corrects it.
Updated Desc
NA
Identifying Info
true
Data Training
Training data is proprietary to Zoom. According to Zoom’s website, “Zoom does not use any customer audio, video, chat, screen sharing, attachments, or other communications-like customer content (such as poll results, whiteboard, and reactions) to train Zoom’s or its third-party artificial intelligence models.”
Data Input
Speech at MOME's rulemaking hearings and agency webinars.
Data Output
Text in a transcript and captions.
Vendor Name
Zoom
Vendor Type
NA
Vendor Desc
Regular updates to the application.
Data 2022
NA
Vendor
NA
Analysis Type
NA
New York Police Department
Year: 2024 • Agency: New York Police Department • Department: NA
Year
2024
Agency
New York Police Department
Department
NA
Tool Name
Evolv Express Weapons Detection System
Date First Use
2024/07
Updated
Created in CY2024
Purpose Type
Data management
Computation Type
Classification
Autonomy
NA
Frequency
NA
Population Type
Individuals; Property; Geographic space
Population Type Individual
General population
Population Type Other
NA
Website
NA
Tool Desc
Electromagnetic weapons detection system.
Purpose Desc
Electromagnetic weapons detection devices emit ultra-low frequency, electromagnetic pulses (similar to those used in retail loss prevention) that pass through objects moving through the system. Sensors process the relayed information and the system uses this data to determine if it detects a potential firearm. The system is equipped with video cameras that are part of a real-time image-aided alert system that will indicate the presence of a firearm to monitoring personnel.
Updated Desc
NA
Identifying Info
true
Data Training
Training data is proprietary to the vendor.
Data Input
Individuals walk through two towers which emit ultra-low frequency, electromagnetic pulses that pass through objects on the individual’s person. Sensors process the information generated and the system uses this data to determine if it detects a potential firearm.
Data Output
If a potential firearm is detected, the system will capture a still image and an approximately three-second video of the individual moving through the system. The system will alert monitoring personnel that a potential firearm has been detected and wirelessly transmit the still image and video to a tablet being monitored by personnel. A cube will appear on both the still image and video clip, indicating the location of the potential firearm being worn or carried by the individual. The location of a cube is discerned by the system based on the electromagnetic data processed by the system sensors.
Vendor Name
EVOLV
Vendor Type
NA
Vendor Desc
Software developed and maintained by vendor.
Data 2022
NA
Vendor
NA
Analysis Type
NA
New York Police Department
Year: 2024 • Agency: New York Police Department • Department: NA
Year
2024
Agency
New York Police Department
Department
NA
Tool Name
Facial Recognition Technology
Date First Use
2011/10
Updated
No
Purpose Type
Data management
Computation Type
Matching
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
General population
Population Type Other
NA
Website
NA
Tool Desc
Tool which may help investigators identify unknown subjects in law enforcement investigations.
Purpose Desc
Facial recognition is a digital technology that NYPD uses to compare images obtained during investigations with lawfully possessed arrest and parole photos. The tool analyzes an uploaded image, known as a probe image, and searches and compares against the image repository. The purpose of the tool is to enhance law enforcement’s ability to investigate criminal activity as well as identify deceased persons and missing persons. When used in combination with human analysis and additional investigation, facial recognition technology is a valuable tool in solving crimes and increasing public safety.
Updated Desc
NA
Identifying Info
true
Data Training
Training data is proprietary to the vendor.
Data Input
If NYPD investigators obtain a still image depicting a face of an unknown individual during an investigation, the image can be submitted for facial recognition analysis in accordance with NYPD facial recognition policy. Known as a probe image, NYPD facial recognition software compares the image to a controlled and limited group of lawfully obtained photos called the photo repository.
Data Output
The facial recognition software will generate a pool of possible match candidates for review by trained Facial Identification Section investigators.
Vendor Name
Dataworks
Vendor Type
NA
Vendor Desc
Software developed and maintained by vendor.
Data 2022
NA
Vendor
NA
Analysis Type
NA
New York Police Department
Year: 2024 • Agency: New York Police Department • Department: NA
Year
2024
Agency
New York Police Department
Department
NA
Tool Name
Patternizr
Date First Use
2016/12
Updated
Yes
Purpose Type
Data management
Computation Type
Matching
Autonomy
NA
Frequency
NA
Population Type
Individuals; Property; Geographic space; Other
Population Type Individual
NA
Population Type Other
Crime classification
Website
NA
Tool Desc
Aids crime analysis in detection of potential crime patterns.
Purpose Desc
Patternizr compares features of crimes and finds ones that are similar and may be part of a crime pattern. Analysts will look at the candidate crimes and suggest the formation of crime patterns to a pattern identification module. If a pattern is formed, detectives often consolidate the investigative efforts (e.g., one detective investigates all the crimes in the pattern.) The report filters non-normal trends into a spreadsheet and displays year-over-year counts of crimes that have non-normal trends. The tool requires a human user to evaluate the output data to see if complaints identified as similar are, in fact, connected to a pattern.
Updated Desc
Routine maintenance.
Identifying Info
false
Data Training
Separate models were trained for each of three different crime types (burglaries, robberies, and grand larcenies). These crime types have a sufficient corpus of prior manually identified patterns for use as training examples. This corpus consists of approximately 10,000 patterns between 2006 and 2015 from each crime type. A portion of this corpus includes complaint records where the same individual was arrested for multiple crimes of the same type within a span of two days.
Data Input
The input data is a candidate crime and its features. A complaint describes details of the crime, including the date and time (which can be a range if the precise time of occurrence is unknown), location, crime subcategory, modus operandi, and suspect information. This information is used to calculate the five types of crime-to-crime similarities used as features by Patternizr: location, date-time, categorical, suspect and unstructured text.
Data Output
Probability that a complaint is connected to a pattern.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
NA
New York Police Department
Year: 2024 • Agency: New York Police Department • Department: NA
Year
2024
Agency
New York Police Department
Department
NA
Tool Name
ShotSpotter
Date First Use
2015/03
Updated
Yes
Purpose Type
Data management
Computation Type
Classification
Autonomy
NA
Frequency
NA
Population Type
Geographic space
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Provides acoustic gunshot detection to assist with emergency call response.
Purpose Desc
Provides acoustic gunshot detection to assist with emergency call response. The tool supports patrol operations in alerting units to potential gunfire and enhances investigations involving firearms.
Updated Desc
Routine maintenance.
Identifying Info
false
Data Training
Training data is proprietary to the vendor.
Data Input
Specialized software analyzes audio signals for potential gunshots.
Data Output
The tool determines the location of the sound source, and once classified as potential gunfire sends the incident to acoustic experts for additional analysis. Notifications are sent for confirmed gunfire. ShotSpotter activations may result in evidence collection that can enhance case investigations. Problematic locations identified through alerts may require additional resource deployment and/or investigations.
Vendor Name
Shotspotter
Vendor Type
NA
Vendor Desc
Software developed and maintained by vendor.
Data 2022
NA
Vendor
NA
Analysis Type
NA
NYC Public Schools
Year: 2024 • Agency: NYC Public Schools • Department: Division of Instructional and Information Technology
Year
2024
Agency
NYC Public Schools
Department
Division of Instructional and Information Technology
Tool Name
Algebra Teaching Assistant
Date First Use
2023/05
Updated
No
Purpose Type
Information presentation
Computation Type
Data generation
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
Teachers, students
Population Type Other
NA
Website
NA
Tool Desc
The Algebra Teaching Assistant uses the Division of Instructional and Information Technology (DIIT) AI platform to accesses specific algebra-focused content to provide responses to prompts related to algebra.
Purpose Desc
The tool is used to generate responses to prompts entered by a student or teacher, requesting the generative AI tool to compose a text response to a text input.
Updated Desc
NA
Identifying Info
true
Data Training
The large language model (LLM) has been trained exclusively on curriculum from Illustrative Math.
Data Input
Prompts provided by the users of the system.
Data Output
The output data for the Teaching Assistant is the response generated by specifically developed LLM using the Illustrative Math curriculum.
Vendor Name
Microsoft
Vendor Type
NA
Vendor Desc
Microsoft provided technical guidance for their emerging generative AI technology and built some small modules of code for the specific Teaching Assistant use cases.
Data 2022
NA
Vendor
NA
Analysis Type
NA
NYC Public Schools
Year: 2024 • Agency: NYC Public Schools • Department: Division of Instructional and Information Technology
Year
2024
Agency
NYC Public Schools
Department
Division of Instructional and Information Technology
Tool Name
Annual Professional Performance Review Measures of Student Learning (MOSL) Growth Model
Date First Use
2013/09
Updated
No
Purpose Type
Performance evaluation
Computation Type
Scoring
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
Teachers
Population Type Other
NA
Website
NA
Tool Desc
The growth model uses a variety of student-level (assessment scores, English language learner, disability, and economic disadvantage indicators), classroom-level (e.g. percent students with disabilities), and school-level data (e.g. percent English language learners, percent students with disability, average prior achievement, school type) to estimate/predict a student’s score on one of many possible course-culminating assessments. These predicted scores are either used to identify “peer groups” of students, from which student growth percentiles (SGPs) are determined, or compared to actual scores to determine student credit values. These units (SGPs or credit values) are then weight-averaged to generate an educator-level result - the Measures of Student Learning (MOSL) rating. The MOSL rating is combined with the Measures of Teaching/Leadership Practice (MOTP/MOLP) rating to produce an Overall Rating. Per state law 3012-d, annual ratings “shall be a significant factor in HR decisions.” This is often implemented by making ratings a qualifying/disqualifying element in decision-making concerning employment, tenure, salary, and other professional opportunities.
Purpose Desc
In accordance with New York State law and New York State Education Department regulations, NYCPS developed and maintains a “growth model” to produce MOSL ratings for use in annual professional performance reviews for teachers and principals. The MOSL ratings are combined with MOTP/MOLP ratings to produce an annual Overall Rating for each eligible educator.
Updated Desc
NA
Identifying Info
false
Data Training
The growth model process is employed in both retrospective and prospective ways. In the retrospective version, the results are determined entirely within-sample. In the prospective version, the coefficients of the model are estimated on multiple prior years of data.
Data Input
The growth model makes use of three types of data: students’ end-of-year assessment scores, enrollment and attendance records that link students to teachers and schools, and historical academic and demographic information used to identify groups of similar students.
Data Output
The model outputs an estimate of a student’s score on a course-culminating assessment.
Vendor Name
Education Analytics
Vendor Type
NA
Vendor Desc
Education Analytics provides technical assistance and quality assurance for the growth model.
Data 2022
NA
Vendor
NA
Analysis Type
NA
NYC Public Schools
Year: 2024 • Agency: NYC Public Schools • Department: Division of Instructional and Information Technology
Year
2024
Agency
NYC Public Schools
Department
Division of Instructional and Information Technology
Tool Name
Annual Professional Performance Review Measures of Teaching/Leadership Practice (MOTP/MOLP) Calculation
Date First Use
2013/10
Updated
No
Purpose Type
Performance evaluation
Computation Type
Scoring
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
Principals, assistant principals, teachers
Population Type Other
NA
Website
NA
Tool Desc
Throughout a school year, evaluators observe teachers/principals multiple times and use a rubric to provide a numerical rating on one or more rubric components. These rubric component scores are then weight-averaged according to collectively bargained rules to produce a Measure of Teaching/Leadership Practice (MOTP/MOLP) Rating. The MOTP/MOLP rating is combined with the Measures of Student Learning (MOSL) rating to produce an Overall Rating for each eligible educator. Per state education law 3012-d, annual ratings “shall be a significant factor in HR decisions.” This is often implemented by making ratings a qualifying/disqualifying element in decision-making concerning employment, tenure, salary, and other professional opportunities.
Purpose Desc
In accordance with New York State law and New York State Education Department regulations, NYCPS developed and maintains databases and calculation rules to produce MOTP/MOLP ratings for use in annual professional performance reviews for teachers and principals. The MOTP/MOLP ratings are combined with MOSL ratings to produce an annual Overall Rating for each eligible educator.
Updated Desc
NA
Identifying Info
false
Data Training
Pilot data prior to program launch was used to inform the weights assigned to various rubric components. However, the weights are ultimately determined via collective bargaining.
Data Input
Rubric component numerical ratings.
Data Output
The model outputs a score for teachers and principals.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
NA
NYC Public Schools
Year: 2024 • Agency: NYC Public Schools • Department: Division of Instructional and Information Technology
Year
2024
Agency
NYC Public Schools
Department
Division of Instructional and Information Technology
Tool Name
Eureka! Chatbot
Date First Use
2023/08
Updated
No
Purpose Type
Information presentation
Computation Type
Data generation
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NYCPS staff and parents
Population Type Other
NA
Website
NA
Tool Desc
The Azure Cognitive Services technology and chatbot (internally branded as “Eureka!”) was configured and deployed in August 2023 to be the first response to calls to the NYCPS IT Service Desk. It accesses scripts to handle four common reasons for a user to call or contact the service desk - Password Reset, Create a Ticket, Ticket Status, Request for Information. The chatbot accesses pre-defined scripts to respond to user voice or text input. The user’s request is either serviced, completed and closed by the chatbot, or the user is given the option (at any time) to connect to a live agent.
Purpose Desc
The tool is used to respond to common IT service desk requests - Password Reset, Create a Ticket, Ticket Status, Request for Information. Users can access the tool by phone, by computer through the NYCPS Support Hub application, and from links from Microsoft Teams and other NYCPS systems, such as TeachHub.
Updated Desc
NA
Identifying Info
true
Data Training
Pre-defined scripts designed to respond to four common requests to the IT Service Desk.
Data Input
A voice call or text-based chat session initiated by a user and responded to by the Eureka! chatbot before being handled by a human Service Desk agent.
Data Output
The chatbot generates responses to user-entered prompts based on the training data or forwards the inquiry to a human Service Desk agent. Since its launch in August 2022, the chatbot has handled an average of 1,500 calls and 300 web-based inquiries each day. Approximately 30 percent of the voice calls and 10 percent of the web-based queries have been handled completely by Eureka! without being forwarded to a human Service Desk agent.
Vendor Name
Nagarro and Microsoft
Vendor Type
NA
Vendor Desc
Developed by an IT services vendor (Nagarro) using Microsoft Cognitive services.
Data 2022
NA
Vendor
NA
Analysis Type
NA
NYC Public Schools
Year: 2024 • Agency: NYC Public Schools • Department: Division of Instructional and Information Technology
Year
2024
Agency
NYC Public Schools
Department
Division of Instructional and Information Technology
Tool Name
MySchools - Match
Date First Use
2018/08
Updated
No
Purpose Type
Resource allocation
Computation Type
Matching
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
Students
Population Type Other
NA
Website
NA
Tool Desc
The tool utilizes the Gale-Shapley deferred acceptance algorithm to match applicants to schools. This algorithm has been in existence for many years, used internationally for various purposes. Perhaps most common is its use in the National Resident Matching Program for medical school students.

Deferred acceptance works as an iterative series of steps: students and programs are tentatively matched in each step, but nothing is finalized until the algorithm terminates (hence the deferred).

1. Each student “proposes” to their first choice;
- Programs assign seats to students one at a time;
- When all seats are filled, programs may reject previously accepted students in favor of new applications from students they prefer (e.g., students with a better lottery number);
- Remaining students are rejected;
2. Students rejected in the last step “propose” to the next choice on their list;
3. The algorithm terminates when all students are matched or have proposed to all the programs they listed.
Purpose Desc
MySchools is an application used to house online school directories, collect application choices, and run the admissions matching algorithm that is used for all centralized admissions processes (3-K, pre-K, Gifted & Talented, middle school, and high school). The tool encompasses a family-facing portal, a school-facing portal, and an administrative portal.
Updated Desc
NA
Identifying Info
true
Data Training
The algorithm was already widely recognized for its advantages prior to adoption in New York City. NYCPS consulted with a team of researchers at Massachusetts Institute of Technology who had been closely involved in its initial creation when we adopted it.
Data Input
Student biographical information (e.g., home address, poverty status, home language), student academic information (e.g., course grades, state test scores), and student school records (e.g., sending school).
Data Output
The algorithm outputs a school match for each student.
Vendor Name
Blenderbox
Vendor Type
NA
Vendor Desc
We have a five-year contract with the agency Blenderbox who designed the application and implemented the algorithmic matching functionality. The work is meant to transition to be run in-house, by the Division of Instructional and Information Technology (DIIT) within NYCPS, by the end of the contract. The team at DIIT has already begun to takeover maintenance and development of the tool.
Data 2022
NA
Vendor
NA
Analysis Type
NA
NYC Public Schools
Year: 2024 • Agency: NYC Public Schools • Department: Division of Instructional and Information Technology
Year
2024
Agency
NYC Public Schools
Department
Division of Instructional and Information Technology
Tool Name
MySchools - Probability of Acceptance
Date First Use
2024/09
Updated
Created in CY2024
Purpose Type
Information presentation
Computation Type
Scoring
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
Students
Population Type Other
NA
Website
NA
Tool Desc
This feature, added to MySchools in 2024, determines a “Probability of Acceptance at a specific school” for a future high school student. This is calculated and displayed as a student is selecting schools to apply to in the MySchools application.
Purpose Desc
Information is presented to students and parents to help them decide what high schools to apply to.
Updated Desc
NA
Identifying Info
true
Data Training
Model was developed by researchers affiliated with the Massachusetts Institute of Technology (MIT) and trained on information about NYC public high schools. Students will see an icon indicating whether they have a “high,” “medium,” or “low” chance of receiving an offer, based on the applicant’s admissions characteristics like district or borough, grades, priority group, and the school’s admissions method, such as whether the admission is open or screened.
Data Input
Student high school selections and student records.
Data Output
Probability of acceptance for the student to a specific high school, indicated as “high”, “medium” or “low”.
Vendor Name
Researchers from MIT
Vendor Type
NA
Vendor Desc
MIT developed the tool and the Division of Instructional and Information Technology integrated it into the MySchools system.
Data 2022
NA
Vendor
NA
Analysis Type
NA
NYC Public Schools
Year: 2024 • Agency: NYC Public Schools • Department: Division of Instructional and Information Technology
Year
2024
Agency
NYC Public Schools
Department
Division of Instructional and Information Technology
Tool Name
Open Gen AI and Teaching Assistant Tool
Date First Use
2023/05
Updated
No
Purpose Type
Information presentation
Computation Type
Data generation
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
Teachers, students
Population Type Other
NA
Website
NA
Tool Desc
The generative AI system using large language models was a system custom-built by the Division of Instructional and Information Technology using advanced Microsoft technologies to create a set of generative AI tools. “Open Gen AI” accesses a large language model (currently OpenAI’s GPT 3.5) to provide responses to a broad range of prompts.
Purpose Desc
The tool is used to generate responses to prompts entered by a student or teacher, requesting the generative AI tool to compose a text response to a text input.
Updated Desc
NA
Identifying Info
true
Data Training
ChatGPT training data is proprietary to OpenAI.
Data Input
Prompts provided by the users of the system.
Data Output
The output data for the Open Gen AI tool is the response generated by the ChatGPT large language model.
Vendor Name
Microsoft
Vendor Type
NA
Vendor Desc
Microsoft provided technical guidance for their emerging generative AI technology and built some small module of code for the specific NYCPS Gen AI and Teaching Assistant use cases.
Data 2022
NA
Vendor
NA
Analysis Type
NA
Office of Chief Medical Examiner
Year: 2024 • Agency: Office of Chief Medical Examiner • Department: Department of Forensic Biology
Year
2024
Agency
Office of Chief Medical Examiner
Department
Department of Forensic Biology
Tool Name
STRMix
Date First Use
2017/01
Updated
No
Purpose Type
Data management
Computation Type
Not specified
Autonomy
NA
Frequency
NA
Population Type
Individuals; Biological sample
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
STRmix™ combines sophisticated biological modeling and standard mathematical processes to interpret a wide range of complex DNA profiles. Using well-established statistical methods, the software builds millions of conceptual DNA profiles. It grades them against the evidential sample, finding the combinations that best explain the profile. A range of likelihood ratio options are provided for subsequent comparisons to reference profiles. Using a Markov Chain Monte Carlo engine, STRmix™ models any types of allelic and stutter peak heights as well as drop-in and drop-out behavior. It does this rapidly, accessing evidential information previously out of reach with traditional methods. STRmix™ is supported by comprehensive empirical studies with its mathematics readily accessible to DNA analysts, so results are easily explained in court.
Purpose Desc
STRMix is a probabilistic genotyping tool that is used to analyze mixtures of DNA profiles to help associate the crime scene evidence to potential victims or suspects of crimes.
Updated Desc
NA
Identifying Info
true
Data Training
Training data was not used in the sense of AI software. OCME performed thousands of tests using the software to validate it for optimum use with our current laboratory standard operating procedures and genetic analyzers.
Data Input
Forensic DNA profiles from crime scenes as well as the DNA profiles from victims and suspects of crimes.
Data Output
The output is a deconvolution of genotype probability distribution that lists all of the accepted genotype sets and their associated weights. These weights can take any value from zero to one.
Vendor Name
NicheVision Forensics, LLC
Vendor Type
NA
Vendor Desc
The software has been developed by New Zealand Crown Institute of Environmental Science and Research with Forensic Science South Australia. The developer assisted OCME in analyzing and interpreting our data during the validation of the software.
Data 2022
NA
Vendor
NA
Analysis Type
NA
Office of Technology and Innovation
Year: 2024 • Agency: Office of Technology and Innovation • Department: Web Operations
Year
2024
Agency
Office of Technology and Innovation
Department
Web Operations
Tool Name
Google Translate
Date First Use
2013/00
Updated
No
Purpose Type
Information presentation
Computation Type
Data generation
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
Users of nyc.gov
Population Type Other
NA
Website
NA
Tool Desc
Google Translate enables machine translation of nyc.gov and subpages into over 100 languages.
Purpose Desc
The Google Translate widget is used to make information on nyc.gov and its subpages more accessible to New Yorkers with limited English proficiency. It translates content into the 10 languages required under the language access law (Local Law 30 of 2017) and over 100 others.
Updated Desc
NA
Identifying Info
true
Data Training
Training data is proprietary to Google.
Data Input
Web content on nyc.gov and its subpages, written in English.
Data Output
Translated web content into the selected language.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
NA
Office of Technology and Innovation
Year: 2024 • Agency: Office of Technology and Innovation • Department: Applications
Year
2024
Agency
Office of Technology and Innovation
Department
Applications
Tool Name
MyCity Chatbot
Date First Use
2023/09
Updated
Yes
Purpose Type
Information presentation
Computation Type
Data generation
Autonomy
NA
Frequency
NA
Population Type
Individuals; Group, organization, or business
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The NYC MyCity chatbot is a beta AI-powered chatbot that provides information and access to services for residents and businesses in New York City.
Purpose Desc
The NYC MyCity chatbot is a beta AI-powered chatbot that provides information and access to services for residents and businesses in New York City. It’s currently focused on two main areas: Business Services and MyCity Basics. The chatbot provides information on starting or operating a business in New York City, answers questions about permits, licenses, regulations, and other business requirements, and connects users with relevant resources and support services. It also offers information on various city services and benefits, and helps users find resources related to childcare, career, and other areas. The chatbot is using Microsoft’s Azure AI technology and OpenAI’s ChatGPT 4-o large language model (LLM).
Updated Desc
Security updates and LLM upgraded to use ChatGPT 4-o.
Identifying Info
false
Data Training
Training data is proprietary to the vendor.
Data Input
Text queries are input by the user on the MyCity portal.
Data Output
The tool produces text responses with references based on information from Business Services and MyCity Basics.
Vendor Name
Microsoft, EY
Vendor Type
NA
Vendor Desc
Microsoft provides cloud-based ChatGPT services, and EY is the professional services vendor.
Data 2022
NA
Vendor
NA
Analysis Type
NA
Office of Technology and Innovation
Year: 2024 • Agency: Office of Technology and Innovation • Department: NYC311
Year
2024
Agency
Office of Technology and Innovation
Department
NYC311
Tool Name
NYC311 AI Voice Pilot
Date First Use
2024/03
Updated
Created in CY2024
Purpose Type
Information presentation
Computation Type
Data generation
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
NYC311 AI Voice pilot is a large language model (LLM)-powered voice call solution that provided information and access to services for residents, businesses, and visitors.
Purpose Desc
The NYC311 AI Voice was a pilot program that enabled NYC311 to test a generative AI voice application with NYC311 customers, for 21 days. On a daily basis, customers can reach the NYC311 call center, by dialing 311, 212-NEW-YORK, or 211. The pilot was kept on a small scale, available only to customers who dialed 211. The NYC311 AI Voice provided information to residents, businesses, and visitors on a wide range of inquiries, as well as providing updates on city programs, events, and notifications using NYC311’s content.
Updated Desc
NA
Identifying Info
false
Data Training
Training data is proprietary to the vendor.
Data Input
The customer made voice inquiries when contacting NYC311.
Data Output
NYC311 AI Voice provided the customer with information/responses from the NYC311 Content Application Programming Interface.
Vendor Name
Microsoft, Nuance
Vendor Type
NA
Vendor Desc
Microsoft provided an LLM-powered solution to aid in handling customers who call NYC311 for information and services. The focus was handling voice interactions.
Data 2022
NA
Vendor
NA
Analysis Type
NA
Office of Technology and Innovation
Year: 2024 • Agency: Office of Technology and Innovation • Department: NYC311
Year
2024
Agency
Office of Technology and Innovation
Department
NYC311
Tool Name
Omnichannel Language Translation
Date First Use
2024/01
Updated
Created in CY2024
Purpose Type
Information presentation
Computation Type
Data generation
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
Customers who contact NYC311 via text/SMS
Population Type Other
NA
Website
NA
Tool Desc
The Omnichannel Language Translation tool delivers multi-language capability for the NYC311 text/SMS channel. The tool supports the 10 designated citywide languages to enable NYC311 agents to interact with customers in their language.
Purpose Desc
The algorithmic tool converts the customer’s text inquiry in their chosen language into English, allowing the text agent to understand, research and reply to the inquiry. The tool converts the agent’s English language response to the customer’s chosen language among the 10 designated citywide languages.
Updated Desc
NA
Identifying Info
false
Data Training
The training data is proprietary to Microsoft.
Data Input
Text/SMS inquiries from customers via 311-NYC.
Data Output
Responses to customer inquiries in the language the customer used.
Vendor Name
Microsoft
Vendor Type
NA
Vendor Desc
Omnichannel is part of the Microsoft suite available to OTI as part of the Dynamics customer relationship management platform. Microsoft supported the design, development, and testing of the tool preparation and deployment.
Data 2022
NA
Vendor
NA
Analysis Type
NA
Office of Technology and Innovation
Year: 2024 • Agency: Office of Technology and Innovation • Department: Office of Data Analytics
Year
2024
Agency
Office of Technology and Innovation
Department
Office of Data Analytics
Tool Name
Zoom Automated Captions
Date First Use
2021/03
Updated
No
Purpose Type
Information presentation
Computation Type
Data generation
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
Attendees of Open Data Week and Open Data Ambassadors meetings who have spoken during the meeting or whose names have been mentioned during the meeting
Population Type Other
NA
Website
NA
Tool Desc
Creates virtual closed captioning/live transcription during Zoom meetings.
Purpose Desc
Captions are provided to attendees of Zoom meetings held by the NYC Open Data team at the Office of Data Analytics in conjunction with the civic tech non-profit BetaNYC under the Open Data Week and Open Data Ambassador initiatives. The full transcription of the event is then added to the meeting recordings, which are uploaded on YouTube. The purpose of the captions, both for the live event and the recording, is to improve meeting accessibility.
Updated Desc
NA
Identifying Info
true
Data Training
Training data is proprietary to Zoom.
Data Input
Live audio from Zoom meeting.
Data Output
VTT format file including captions/transcript of meeting.
Vendor Name
BetaNYC
Vendor Type
NA
Vendor Desc
BetaNYC is our collaborator on the Open Data Week and Open Data Ambassador initiatives. They own and operate the Zoom account that is used for meetings under these initiatives, have access to the transcription files, and use these when editing and uploading video recordings onto YouTube.
Data 2022
NA
Vendor
NA
Analysis Type
NA
Administration for Children's Services
Year: 2023 • Agency: Administration for Children's Services • Department: NA
Year
2023
Agency
Administration for Children's Services
Department
NA
Tool Name
Accelerated Safety Analysis Protocol (ASAP) Tool
Date First Use
2018/05
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Predictions of Severe Harm (identifying likelihood of substantiated allegations of physical or sex abuse within the next 24 months) are based on a machine learning methodology and are calculated for all children involved in active investigations early in the investigation (day 10). An investigation is assigned a numeric likelihood of this outcome based on the child in the case with the highest likelihood. The ACS Quality Assurance unit in the Division of Child Protection reviews about 3,000 active investigations annually, selecting those with the highest likelihood of severe harm.
Purpose Desc
The Quality Assurance Unit in the Division of Child Protection at ACS has the capacity to review about 3,000 investigation cases out of about 50,000 investigations annually. ACS developed this predictive model to support the selection of cases for Quality Assurance review. Open investigations involving children with the greatest likelihood to experience future severe harm – substantiated allegations of physical or sex abuse in the following 24 months – are selected for review. The tool does not support decisions about services or interventions for individuals or families involved with ACS, beyond the selection of the case for this additional Quality Assurance review.

If the Quality Assurance review team identifies gaps in routine, required documentation or practice, the team speaks with the field office conducting the investigation and follows up to make certain these gaps have been addressed. Scores are not shared with staff in the Quality Assurance unit or the investigative unit. The model only supports the decision about which investigations are prioritized for review by the Quality Assurance (QA) unit.
Updated Desc
NA
Identifying Info
true
Data Training
ACS trained the model on ACS historic administrative data about closed investigations from April 2014 to April 2016. The training set included about 142,026 observations. The model was tested on closed investigations from April 2016 to April 2017 with 53,477 observations.
Data Input
Predictions are based on administrative data about prior and current child welfare involvement including investigations triggered by a New York State Central Register (SCR) call and time spent in foster care. Only ACS administrative data are used in the model.
Data Output
Rank ordered list of open investigation cases involving children with the highest likelihood to experience future severe harm, defined as substantiated allegations of physical or sex abuse in the following 24 months to be reviewed by a special QA Review Team.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Administration for Children's Services
Year: 2023 • Agency: Administration for Children's Services • Department: NA
Year
2023
Agency
Administration for Children's Services
Department
NA
Tool Name
Housing Prioritization
Date First Use
2023/04
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The city has allocated 100 housing vouchers to families receiving ACS Prevention Services. The shelter application model identifies the likelihood of a family in prevention services applying for homeless shelter within 12 months beyond the current prevention case. The model uses a machine learning methodology and is calculated for all children in a prevention case. ACS Prevention Services reaches out to the Service Providers assisting the families with the highest risk for applying for shelter.
Purpose Desc
The model estimates a risk score for a child receiving prevention services whose family will apply and be eligible for a homeless shelter within 12 months from the start of service (day 14 of the prevention case).

The model helps predict the risk of application for homeless shelters among families receiving prevention services. With a limited number of vouchers available, the risk model helps ACS prioritize housing assistance for those families at greatest risk of becoming homeless.

The Service Provider meets with the family to conduct a qualitative assessment of the family’s housing needs and vouchers are offered based on their findings.
Updated Desc
NA
Identifying Info
true
Data Training
ACS trained the model on ACS historic administrative data regarding preventive services started between 2014 and 2020. An 80/20 split of data to train on 80 percent and test on 20 percent ensuring that no family appears in both sets. The training set contains 140,242 observations between Jan 2014 and December 2020. The test set consisted of 34,508 observations between Jan 2014 and December 2020.
Data Input
Predictions are based on administrative data about prior and current child welfare involvement at the start of a case. This includes SCR investigations and time spent in foster care. Only ACS administrative data are used in the model.
Data Output
Rank ordered list of open prevention cases involving children whose families have the highest likelihood of applying for a homeless shelter within 12 months of starting a prevention service.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Administration for Children's Services
Year: 2023 • Agency: Administration for Children's Services • Department: NA
Year
2023
Agency
Administration for Children's Services
Department
NA
Tool Name
Prevention Score Card
Date First Use
2021/09
Updated
Yes
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Predictions of Repeat Maltreatment (identifying the likelihood of being involved in a future indicated investigation within the next 24 months at the start of service) are based on a machine learning methodology and are calculated for all children receiving prevention services from ACS prevention service providers.
Purpose Desc
The Repeat Maltreatment model is used to make predictions on day 10 from the start of the prevention case to assess the risk of the family at the beginning of the service. A prevention case is assigned a numeric likelihood of an indicated investigation based on a New York State Central Register (SCR) within 24 months from the start of a prevention service.

The prevention providers are assessed for their performance based on the service needs/risk levels of the families they’ve served during the previous fiscal year.

The programs were sorted and ranked based on their average risk, and then divided into four quartiles by rank order: the top 25 percent of programs are classified as the Very High-Risk Cohort, the next 25 percent of programs as the High-Risk Cohort, the next 25 percent as the Medium-Risk Cohort, and the lowest 25 percent as the Low-Risk Cohort. Assignment to a cohort is not a way of performance assessment of the program but to group prevention service providers for fair comparisons based on the risk level of families they served.
Updated Desc
In August 2023, ACS retrained the model on recent ACS historic administrative data about closed investigations with an 80/20 split in data from July 2009 to June 2018. The training set included 80 percent (about 338, 467) observations. The model was tested on 20 percent of closed investigations from July 2009 to June 2018 with 84,494 observations.
Identifying Info
true
Data Training
ACS trained the model on ACS historic administrative data about closed investigations from July 2009 to June 2016. Training set included about 158,787 observations. The model was tested on closed investigations from July 2016 to June 2018 with 46,969 observations.
Data Input
Predictions are based on administrative data about prior and current child welfare involvement at the start of a case. This includes SCR investigations and time spent in foster care. Only ACS administrative data are used in the model.
Data Output
The model is used for generating a scorecard of prevention service providers by categorizing prevention programs based on the average risk profile of the cases they served during the assessment year. These groupings of program cohorts provide context for understanding the Scorecard, as it allows for performance comparison of programs that accepted and served families with similar risk profiles.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Administration for Children's Services
Year: 2023 • Agency: Administration for Children's Services • Department: NA
Year
2023
Agency
Administration for Children's Services
Department
NA
Tool Name
Service Termination Conference (STC)
Date First Use
2017/07
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Predictions of Repeat Maltreatment (identifying the likelihood of being involved in a future indicated investigation within the following 24 months towards the end of service) are based on a machine learning methodology and are calculated for all children receiving prevention services from ACS prevention service providers. The prediction is made with the assumption that the case is closed the day the model is run. A prevention case is assigned a numeric likelihood of an indicated investigation based on a New York State Central Register (SCR) within 24 months from the end of a prevention service.
Purpose Desc
The Repeat Maltreatment model was known as the Service Termination Conference (STC) model and was used by Preventive Services managers to identify whether or not a case should have ACS or provider-agency facilitation at the service termination conference. The model was initially used to prioritize ACS facilitation of prevention termination conferences so that ACS could be certain that services had been provided in these cases. When a family is ready to exit ACS prevention services, an end-of-services conference is required (known as a “Service Termination Conference”). However, Service Termination Conferences are generally no longer facilitated by ACS and this prioritization is no longer necessary.

In November 2022, a pilot was initiated with the STC list to recommend low-risk families to prevention service providers in case they were interested in discontinuing service. This pilot was stopped in January 2023.
Updated Desc
NA
Identifying Info
true
Data Training
ACS trained the model on ACS historic administrative data about closed investigations from July 2009 to December 2015. Training set included about 130,982 observations. The model was tested on closed investigations from January 2016 to December 2018 with 48,771 observations.
Data Input
Predictions are based on administrative data about prior and current child welfare involvement at the end of a prevention service case. It includes SCR investigations and time spent in foster care. Only ACS administrative data are used in the model.
Data Output
As of September 2021, ACS is no longer required to facilitate prevention termination conferences (all conferences are facilitated by the contract prevention programs) and the STC model is no longer being used for this purpose.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Administration for Children's Services
Year: 2023 • Agency: Administration for Children's Services • Department: NA
Year
2023
Agency
Administration for Children's Services
Department
NA
Tool Name
Un-Involvement Model
Date First Use
2023/05
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Predictions of “Un-Involvement” (identifying the likelihood of no future involvement with ACS within 24 months beyond the current investigation) are based on a machine learning methodology and are calculated for all children in an ongoing investigation. This future engagement for families may be in the form of ACS prevention services, court-ordered supervision, or foster care services. There are two models run on day 10 and day 40 to ensure that a case that was recommended for early closure on day 10 is still eligible and recommended on day 40, right before actually closing the case.
Purpose Desc
The model helps Child Protection managers identify low-risk cases that are likely not to require further ACS involvement beyond the “current” investigation and therefore could be considered for early closure, sooner than the typical 60 days.

The model generates initial risk predictions on the 10th day of a new investigation to identify low-risk cases. Subsequently, on the 40th day of the investigation, the model re-evaluates these cases with a new set of risk predictions to determine if they continue to be classified as low-risk. If new information collected during the intervening 30 days suggests that a case is no longer eligible or is no longer considered low-risk, the recommendation to close the case will be revised.
Updated Desc
NA
Identifying Info
true
Data Training
ACS trained the model on ACS historic administrative data about closed investigations from 2012 to 2017. An 80/20 split of data to train on 80 percent and test on 20 percent ensuring that no family appears in both sets. The training set contains 381,649 children from 183,516 investigations ending between Jan 2012 and December 2017. The test set consisted of 101,369 children from 48,794 investigations ending between Jan 2012 and December 2017.
Data Input
Predictions are based on administrative data about prior and current child welfare involvement at the start of a case. This includes SCR investigations and time spent in foster care. Only ACS administrative data are used in the model.
Data Output
A recommendation and description are displayed on both day 10 and day 40 of all open cases via a reporting platform, viewable to only deputy directors. Upon discussion with the deputy directors on day 40, the case worker makes a determination to close the case or not. The case workers are not aware that the recommendation is generated by a machine learning model, to not bias the decision-making process. Alternatively, the deputy directors use the list to inform their workload planning conversations with managers and staff, when caseloads are high.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Department of Consumer and Worker Protection
Year: 2023 • Agency: Department of Consumer and Worker Protection • Department: NA
Year
2023
Agency
Department of Consumer and Worker Protection
Department
NA
Tool Name
Route Automation
Date First Use
2020/07
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Group, organization, or business
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Inspection Supervisor selects an inspector, enters a date and the number of businesses to be inspected, and the geographic area to be considered. The system identifies businesses in the selected area and assigns them to the route based on inspection priority until the number of businesses entered has been reached. Then the tool runs a Simulated Annealing Algorithm to optimize the order businesses appear on the route based on proximity and method of travel.
Purpose Desc
DCWP inspectors conduct inspections based on a route, or list of businesses to be inspected on a specific day, which must be pre-approved by their supervisor. The Route Automation tool generates a route for an inspector on a specific date based on configuration variables and geographic area. All routes generated by the tool still require supervisor review and approval.
Updated Desc
NA
Identifying Info
false
Data Training
The tool is not trained in an AI or Machine Learning sense. The tool makes decisions based on configuration tables, businesses and their licenses (if any), and inspection and violation history, and uses a Simulated Annealing Algorithm to optimize the order in which businesses appear on a route.
Data Input
Inspection date, business category and address, licenses held (if any), last inspection date and type, violation history (if any), date of inspection request or license application/renewal (if applicable), Inspection Unit (of the inspector), and geographic area.
Data Output
An ordered list of businesses to be inspected on a given day by a given inspector.
Vendor Name
PruTech
Vendor Type
NA
Vendor Desc
The tool was designed and built by PruTech, an outside vendor contracted to design and build DCWP’s Automated Inspection Management System (AIMS) and its accompanying Mobile Enforcement platform. The tool is part of the AIMS system.
Data 2022
NA
Vendor
NA
Analysis Type
Optimization
Department of Environmental Protection
Year: 2023 • Agency: Department of Environmental Protection • Department: NA
Year
2023
Agency
Department of Environmental Protection
Department
NA
Tool Name
Idling Complaints Program
Date First Use
2022/08
Updated
Yes
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
A contractor helped create an AI tool that analyzes the audio and visual aspects of pictures and videos submitted by citizens of alleged car idling complaint occurrences that are in violation of New York City air pollution laws.
Purpose Desc
The analysis from the tool makes a recommendation to staff reviewers whether the submitted evidence support an occurrence of car idling in violation of New York City laws. The tool also provides a level of confidence in its recommendation. The tool does not make the review decision in the Idling Complaints system. It is still entirely up to the staff to decide whether to take the tool’s recommendation or not.
Updated Desc
Retraining model with newer data, modified the description text generated to be more useful for reviewer staff.
Identifying Info
false
Data Training
Videos and pictures of cars idling submitted by citizens, along with staff decisions on whether the picture/video constituted as an idling violation.
Data Input
Videos and pictures submitted by citizens through our web portal.
Data Output
Recommendation, confidence level, description of its decision from the tool.
Vendor Name
Acuvate
Vendor Type
NA
Vendor Desc
Acuvate developed the AI tool that performs the automated analysis of the submitted evidence.
Data 2022
NA
Vendor
NA
Analysis Type
Computer vision
Department of Health and Mental Hygiene
Year: 2023 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2023
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
Bionumerics
Date First Use
2017/09
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals; Other
Population Type Individual
NA
Population Type Other
Sequence data can belong to any species
Website
NA
Tool Desc
A suite of tools used to align and analyze bacterial genomes.
Purpose Desc
BioNumerics is used to 1) re-assemble the bacterial genome (since the sequencing process involves fragmenting the bacterial DNA and then amplifying it into millions of pieces) 2) identify the genus; species; and serotype of the bacterial isolate 3) perform quality control checks to ensure the sequence meets certain quality standards 4) perform whole genome multi-locus sequence typing (wgMLST; a technique used to type bacteria based on their genetic code) 5) perform cluster analysis for cases related to one another based upon case definitions recommended by the Centers for Disease Control and Prevention (CDC). This information is then communicated to partners including foodborne epidemiologists at the Bureau of Communicable Disease; who investigate all reported cases of foodborne disease; with those investigations potentially resulting in restaurant inspections; closures; and food recalls.
Updated Desc
NA
Identifying Info
false
Data Training
N/A
Data Input
Fastq.
Data Output
txt; Excel.
Vendor Name
Bionumerics
Vendor Type
NA
Vendor Desc
Developed and maintains the tool.
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling; Matching
Department of Health and Mental Hygiene
Year: 2023 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2023
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
Bowtie2
Date First Use
2022/06
Updated
Yes
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals; Other
Population Type Individual
NA
Population Type Other
Sequence data can belong to any species
Website
NA
Tool Desc
Aligns sequencing data to a reference sequence. Bowtie2 aligns sequencing data to a reference using Burrows-Wheeler transformations. It is geared to use with Illumina sequencing data.
Purpose Desc
Bowtie2 is an intermediate step in the workflow to analyze COVID variants in wastewater.
Updated Desc
Updated variable weights/bug fixes/dependencies.
Identifying Info
false
Data Training
N/A
Data Input
Sequence reads (fastq) for single or paired-end runs (sequence reads can be considered strings).
Data Output
Aligned reads in SAM format
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling; Matching