Algorithmic Tools Compliance Report

"This is an annual report on algorithmic tools used by City agencies, collected under Local Law 35 of 2022 (LL 35). It includes descriptions of the tool's use and purpose, datasets used, and vendor involvement. The report is also published as a PDF on the OTI website: https://www.nyc.gov/content/oti/pages/reports.

The full text of LL 35 is available online: https://legistar.council.nyc.gov/LegislationDetail.aspx?ID=4265421&GUID=FBA29B34-9266-4B52-B438-A772D81B1CB5

An "algorithmic tool" is defined by the law as: "Any technology or computerized process that is derived from machine learning, artificial intelligence, predictive analytics, or other similar methods of data analysis, that is used to make or assist in making decisions about and implementing policies that materially impact the rights, liberties, benefits, safety or interests of the public, including their access to available city services and resources for which they may be eligible. Such term includes, but is not limited to tools that analyze datasets to generate risk scores, make predictions about behavior, or develop classifications or categories that determine what resources are allocated to particular groups or individuals, but does not include tools used for basic computerized processes, such as calculators, spellcheck tools, autocorrect functions, spreadsheets, electronic communications, or any tool that relates only to internal management affairs such as ordering office supplies or processing payments, and does not materially affect the rights, liberties, benefits, safety or interests of the public."

City Government Office of Technology and Innovation (OTI) Dataset jaw4-yuem 27 fields
DOWNLOAD CSV
Dataset fields
Showing 50 real records
Department of Health and Mental Hygiene
Year: 2023 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2023
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
BWA
Date First Use
2017/07
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals; Other
Population Type Individual
NA
Population Type Other
Sequence data can belong to any species
Website
NA
Tool Desc
Aligns sequencing data to a reference sequence.
Purpose Desc
Burrows-Wheeler Aligner (BWA) is aligning sequence data to reference using Burrows-Wheeler transformations. This tool is optimal for low-divergent genomic data and short read data; such as Illumina sequence data. This tool is used to predict the order in which the fragments generated by sequencers are pieced together to form a complete genomic sequence data. This tool is used for Legionella and PulseNet sequencing analyses.
Updated Desc
NA
Identifying Info
false
Data Training
N/A
Data Input
Sequence reads (fastq) for single or paired-end runs (sequence reads can be considered strings).
Data Output
Aligned reads in SAM format
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling; Matching
Department of Health and Mental Hygiene
Year: 2023 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2023
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
ChoiceMaker (CM)
Date First Use
2003/06
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
ChoiceMaker (CM) is a record-matching tool that identifies duplicate records belonging to the same individual.
Purpose Desc
CM is used by BOI and Healthy Homes to identify duplicate immunization and lead records. The outputs produced by CM are used in ongoing manual and automated deduplication processes (record merging).
Updated Desc
NA
Identifying Info
true
Data Training
N/A
Data Input
CM uses demographic data (e.g.; names; date of birth; address; identifiers) and health event data (e.g.; date and type of event) from BOI’s Citywide Immunization Registry (CIR) and Healthy Homes’ LeadQuest registry in its evaluation.
Data Output
The program outputs a series of record pairs and a match probability for each pair.
Vendor Name
HLN Consulting
Vendor Type
NA
Vendor Desc
A vendor was involved in the development of the program initially. CM is now available as an open-source program. The DOHMH implementation is maintained by HLN Consulting.
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling; Optimization; Matching
Department of Health and Mental Hygiene
Year: 2023 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2023
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
GATK
Date First Use
2017/10
Updated
Yes
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals; Other
Population Type Individual
NA
Population Type Other
Sequence data can belong to any species
Website
NA
Tool Desc
A suite of tools for variant calling and filtering after sequence alignment. It uses naive Bayesian to qualify aligned bases as sequence or erroneous data; which would be excluded from the final genomic sequence.
Purpose Desc
Used to identify mutations and call upon differences from the reference; which is used to generate the predicted complete sequence.
Updated Desc
Updated variable weights/bug fixes/dependencies.
Identifying Info
false
Data Training
Sets of known variant sites.
Data Input
Fasta; uBam; SAM/BAM/CRAM; VCF.
Data Output
Bam; txt; vcf.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Department of Health and Mental Hygiene
Year: 2023 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2023
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
Guppy
Date First Use
2020/06
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals; Other
Population Type Individual
NA
Population Type Other
Sequence data can belong to any species
Website
NA
Tool Desc
Converts electric signals to predict a nucleotide and enables filtering of low-quality calls.
Purpose Desc
This is a tool designed specifically for Oxford Nanopore Technology (ONT) data. This is a neural network based basecaller; a tool that determines nucleotide bases of a genetic material; that converts electric signals into strings to represent genomic data. In addition to basecalling; the tool also performs filtering of low-quality reads; a stretch of sequenced genetic material. This is the initial step that converts electric signals to fragments of sequence data; which can then be used for COVID-19 sequencing analysis.
Updated Desc
NA
Identifying Info
false
Data Training
The default models within Guppy are trained on a mixture of native and amplified DNA/RNA; from multiple organisms including plant; animal; bacterial and viral genomes.
Data Input
DNA/RNA strand passing through the nanopore. Raw data is stored as .fast5 files
Data Output
.fast5 files; fastq; or BAM files.
Vendor Name
Oxford Nanopore Technologies
Vendor Type
NA
Vendor Desc
Developed and maintains the tool.
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Department of Health and Mental Hygiene
Year: 2023 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2023
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
ICE - Immunization Calculation Engine
Date First Use
1997/00
Updated
Yes
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The Immunization Calculation Engine (ICE) is an immunization evaluation and forecasting system; whose default immunization schedule supports all routine childhood; adolescent; and adult immunizations based on the recommendations of the Advisory Committee on Immunization Practices (ACIP). ICE is free and open-source available through https://cdsframework.atlassian.net/wiki/spaces/ICE/overview.
Purpose Desc
ICE is used by the Bureau of Immunization to evaluate a patient’s immunization history and generate appropriate immunization recommendations.
Updated Desc
New vaccine groups and recommendations were added.
Identifying Info
true
Data Training
N/A
Data Input
ICE uses demographic data (e.g. date of birth) and vaccination data (e.g.; immunization date; vaccine group and type) in the evaluation process. Data used are stored in the Citywide Immunization Registry (CIR).
Data Output
The program returns recommendations on whether a patient has completed a vaccine series or is due for vaccines.
Vendor Name
HLN Consulting
Vendor Type
NA
Vendor Desc
A vendor was involved in the development of the program and continues to be involved in program enhancements. ICE is also available as an open-source program. The DOHMH implementation is maintained by HLN Consulting.
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Department of Health and Mental Hygiene
Year: 2023 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2023
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
Improving Foodborne Disease Outbreak Detection by Incorporating Complaints Identified in Social Media Data
Date First Use
2016/11
Updated
Yes
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals; Group, organization, or business
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Restaurant associated foodborne disease outbreaks are often identified through complaints received via New York City's 311 non-emergency information system; however not all individuals report to 311. The New York City Department of Health and Mental Hygiene (NYC DOHMH) in collaboration with Columbia University developed a text classifier program which monitors Yelp and Twitter data to identify complaints of foodborne illness which was supported by grants from the Alfred P Sloan Foundation and the National Science Foundation. As of April 2023; the tool no longer uses data from Twitter/X due to API changes.
Purpose Desc
The model uses data from Yelp restaurant reviews and previously used Twitter data that was available on Twitter’s publicly available API. Twitter (X) removed free access to their publicly available API in April 2023; so these data are no longer included in our analyses. The classifiers assign a “sick score” to each Yelp review or tweet indicating the likelihood that the review or tweet pertains to foodborne illness. The sick score is based on whether the review/tweet contains key words indicative of foodborne illness (“e.g. vomit”); the Yelp classifier also incorporates if the review indicates that multiple people became sick and if the review indicates a time between eating at a restaurant and illness onset (incubation period) that is consistent with foodborne illness. Each review and tweet with a sick score greater than or equal to a threshold value are reviewed and annotated by DOHMH foodborne disease epidemiology and environmental health staff to determine if the review/tweet was actually reporting foodborne illness possibly associated with a New York City restaurant; if yes; Yelp messages are sent to Yelp reviewers; requesting that they contact DOHMH; and a Twitter message with a survey link was tweeted back to Twitter users to confirm foodborne illness. Data from annotations are used to improve classifier performance. Foodborne disease complaints identified through Yelp and previously Twitter are combined with foodborne disease complaints reported to 311 to improve efficiency of outbreak detection.
Updated Desc
As of April 2023; the tool no longer uses data from Twitter/X due to API changes.
Identifying Info
false
Data Training
Training data was used in the development of both the Yelp and Twitter classifiers. The training data consisted of restaurant reviews and tweets obtained; respectively; from Yelp and Twitter by Columbia University; a subset of these data were joined with annotations provided by DOHMH staff. The annotations of restaurant reviews focused on the following: 1) if the review indicated foodborne illness; 2) if the incident occurred in the past 30 days; 3) if multiple people were sick and 4) if the incubation period was consistent with foodborne illness. For tweets; the annotations focused on if the tweet was indicating foodborne illness and if the incident occurred in New York City. The training data is periodically updated (with annotations from DOHMH) to improve the classifiers.
Data Input
Yelp reviews of New York City restaurants are pulled from a privately available application programming interface (API) provided by Yelp. Publicly available data from Twitter’s API was used through April 2023.
Data Output
The output data includes a “sick score” that the classifiers assign to each Yelp review or tweet (when that data were being used) indicating the likelihood that the review or tweet pertains to foodborne illness.
Vendor Name
Columbia University
Vendor Type
NA
Vendor Desc
DOHMH staff; including Bureau of Communicable Disease; Office of Environmental Investigations; and Division of Informatics and Information Technology & Telecommunications and Columbia University are involved in making decisions about the tool. Columbia University Department of Computer Science professors and doctoral students maintain the classifier. The project was previously funded by the Alfred P Sloan Grant; for which The Fund for Public Health in New York provided administrative support and grant management to DOHMH. This support and management ended at the completion of the grant in 2021.
Data 2022
NA
Vendor
NA
Analysis Type
Speech and language processing
Department of Health and Mental Hygiene
Year: 2023 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2023
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
IQTREE
Date First Use
2020/05
Updated
Yes
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals; Other
Population Type Individual
NA
Population Type Other
Sequence data can belong to any species
Website
NA
Tool Desc
IQTREE uses maximum-likelihood regression to create phylogenetic trees from genomes.
Purpose Desc
Produced phylogenetic trees are used to help rule in or out outbreaks of COVID or other organisms.
Updated Desc
Updated variable weights/bug fixes/dependencies.
Identifying Info
false
Data Training
N/A
Data Input
FASTA; NEXUS; CLUSTALW; PHYLIP.
Data Output
Readable report; ML tree in NEWICH format; log file for entire run.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling; Matching
Department of Health and Mental Hygiene
Year: 2023 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2023
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
kSNP3
Date First Use
2022/03
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals; Other
Population Type Individual
NA
Population Type Other
Sequence data can belong to any species
Website
NA
Tool Desc
kSNP3 can use multiple algorithms (maximum-likelihood; parsimony; neighbor-joining) to infer phylogenetic trees from genomes.
Purpose Desc
Produced phylogenetic trees are used to help rule in or out outbreaks of bacteria.
Updated Desc
NA
Identifying Info
false
Data Training
N/A
Data Input
Fasta
Data Output
ML tree in NEWICH format; log & configuration files; Fasta file
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Department of Health and Mental Hygiene
Year: 2023 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2023
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
MAFFT
Date First Use
2021/01
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals; Other
Population Type Individual
NA
Population Type Other
Sequence data can belong to any species
Website
NA
Tool Desc
Aligns multiple sequencing data.
Purpose Desc
MAFFT (for Multiple Alignment using Fast Fourier Transform) includes several algorithmic methods; including guided tree; scoring matrices; and sequence alignment algorithms to realign multiple genomic sequencing data. The realignment tool is used to locally re-arrange sequence data to make all sequences comparable but the same genomic coordinates. This is used in all sequencing analysis prior to building a phylogenetic tree or distance tree.
Updated Desc
NA
Identifying Info
false
Data Training
N/A
Data Input
Sequences can be in GCG; FASTA; EMBL (Nucleotide only); GenBank; PIR; NBRF; PHYLIP or UniProtKB/Swiss-Prot (Protein only) format
Data Output
Fasta or Clustalw
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling; Matching
Department of Health and Mental Hygiene
Year: 2023 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2023
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
Minimap2
Date First Use
2020/05
Updated
Yes
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals; Other
Population Type Individual
NA
Population Type Other
Sequence data can belong to any species
Website
NA
Tool Desc
Aligns sequencing data to a reference sequence.
Purpose Desc
Minimap2 uses optimal chaining scores to align sequencing data to reference genomes. This tool is faster and more optimal for long read sequences; such as Oxford Nanopore Technologies (ONT) data. This tool is used to predict the order in which the fragments generated by sequencers are pieced together to form a complete genomic sequence data. This tool is used for COVID-19 and monkeypox virus (MPXV) sequencing analyses.
Updated Desc
Fixed the broken Python package. Updated variable weights/bug fixes.
Identifying Info
false
Data Training
N/A
Data Input
Sequence reads (fastq) for single or paired-end runs (sequence reads can be considered strings).
Data Output
Aligned reads in SAM format
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling; Matching
Department of Health and Mental Hygiene
Year: 2023 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2023
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
Pangolin
Date First Use
2021/07
Updated
Yes
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals; Other
Population Type Individual
NA
Population Type Other
Sequence data can belong to any species
Website
NA
Tool Desc
Assigns lineage names to SARS-CoV-2.
Purpose Desc
Pangolin uses a combination of several methods; including random forest tree; classification methods; and maximum parsimony to assign lineage names to SARS-CoV-2 genomic sequences to bin sequences that are more likely to be similar. This is a tool that designates a name based on a nomenclature for COVID-19 sequence data.
Updated Desc
Updated variable weights/bug fixes/dependencies.
Identifying Info
false
Data Training
Trained on a data set of genomes that have been designated to Pango lineages using whole genome information.
Data Input
Fasta files.
Data Output
.csv fiile with taxon name and lineage assigned.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Department of Health and Mental Hygiene
Year: 2023 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2023
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
PHYLOViZ
Date First Use
2017/10
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals; Other
Population Type Individual
NA
Population Type Other
Sequence data can belong to any species
Website
NA
Tool Desc
For representing the possible evolutionary relationships between strains; PHYLOViZ uses the goeBURST algorithm; a refinement of eBURST algorithm by Feil et al.; and its expansion to generate a complete minimum spanning tree (MST).
Purpose Desc
Used to generate the minimum spanning tree relationships.
Updated Desc
NA
Identifying Info
false
Data Training
N/A
Data Input
txt; NEWICK; FASTA.
Data Output
Minimmum spanning tree.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Optimization
Department of Health and Mental Hygiene
Year: 2023 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2023
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
Spades
Date First Use
2017/10
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals; Other
Population Type Individual
NA
Population Type Other
Sequence data can belong to any species
Website
NA
Tool Desc
Spades uses several algorithms to simplify genomic read data into de Brujin graphs and finds overlaps to assemble genomes.
Purpose Desc
Spades is an intermediate step in the workflows of bacterial analyses.
Updated Desc
NA
Identifying Info
false
Data Training
N/A
Data Input
Fastq.
Data Output
Fastas and other files for corrected reads; scaffolds; contigs; paths in GFA format; fastg assembly graph.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling; Matching
Department of Health and Mental Hygiene
Year: 2023 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2023
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
Vsearch
Date First Use
2022/06
Updated
Yes
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals; Other
Population Type Individual
NA
Population Type Other
Sequence data can belong to any species
Website
NA
Tool Desc
Vsearch uses the Needleman-Wunsch algorithm to merge read pairs and align and dereplicate sequences to detect chimeric genomic sequences.
Purpose Desc
Vsearch is an intermediate step in the workflow to analyze COVID variants in wastewater.
Updated Desc
Updated variable weights/bug fixes/dependencies.
Identifying Info
false
Data Training
N/A
Data Input
Sequence reads (fastq; Fasta) for single or paired-end runs (sequence reads can be considered strings).
Data Output
FASTA; FASTQ; tables; alignments; SAM.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Department of Investigation
Year: 2023 • Agency: Department of Investigation • Department: NA
Year
2023
Agency
Department of Investigation
Department
NA
Tool Name
Facial Recognition Technology
Date First Use
2019/03
Updated
Yes
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The tool analyzes an uploaded image or video and searches and compares it with lawfully possessed images to generate a pool of possible matches. If possible matches are identified, a trained DOI examiner visually analyzes and evaluates potential matches to assess reliability of a match consistent with agency policy and applicable laws. A match serves as an investigative lead for additional investigative steps and does not constitute a positive identification.
Purpose Desc
Facial recognition is a digital technology that DOI uses to analyze uploaded images or videos of people and objects obtained during an investigation by comparison with lawfully possessed images. Facial recognition generates possible matches of an object or individual from this analysis and comparison. The purpose of the tool is to assist DOI investigations of matters within its jurisdiction including fraud and other criminal activity.
Updated Desc
DOI acquired additional facial recognition technology in the past year for the above described purpose. DOI's facial recognition tools receive product updates and data sets on an ongoing basis.
Identifying Info
true
Data Training
Self-trained in system usage.
Data Input
Images.
Data Output
Images.
Vendor Name
Not disclosable
Vendor Type
NA
Vendor Desc
Out-of-the-box products. The vendors provide ongoing technical assistance. Confidentiality agreements are in place with the vendors.
Data 2022
NA
Vendor
NA
Analysis Type
Computer vision; Optimization; Matching
Department of Social Services
Year: 2023 • Agency: Department of Social Services • Department: NA
Year
2023
Agency
Department of Social Services
Department
NA
Tool Name
Homebase Risk Assessment Questionnaire (RAQ)
Date First Use
2012/06
Updated
Yes
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Homebase applicants answer screening questions about their current housing situation, history of disruptive experiences, shelter history, and other domains. Each of the answers is assigned a number of points, and applicants that reach a certain point threshold are eligible for deeper Homebase services, such as financial assistance and case management. Workers are able to override a limited number of model decisions with permission of a supervisor.
Purpose Desc
The Homebase program was created to prevent households from entering the DHS shelter system. Since NYC has a range of antipoverty programs and the number of households entering shelter is small compared to the pool of New Yorkers who have an eviction filing each year, the Agency had to ensure that the households who most needed additional homelessness prevention services were being enrolled in Homebase programs. Research showed that staff were not accurately able to predict who would or would not enter the DHS shelter system and that using a risk assessment would provide a better way to match resources to the families who would benefit the most.
Updated Desc
Fresh training data. Started using on 12/21/2023.
Identifying Info
true
Data Training
The RAQ was developed based on analysis of data on Homebase enrollees from 2004 to 2008, conducted in conjunction with a team of academic researchers, to determine predictive factors for those entering shelter. It was updated in 2023 based on analysis led by NYC DSS researchers of 2013-2016 Homebase data.
Data Input
Factors include, among others: personal characteristics such as age and pregnancy; educational attainment and employment status; housing issues such as eviction, discord, a move in the past year; past and recent experience of homelessness.
Data Output
The tool produces a score that is used to assess eligibility for full versus brief Homebase services.
Vendor Name
Multiple researchers
Vendor Type
NA
Vendor Desc
DHS contracted with researchers to evaluate years of Homebase administrative data to develop a risk assessment. The DSS research team then led an updated analysis that led to tool revisions. The published research papers are listed below: https://ajph.aphapublications.org/doi/10.2105/AJPH.2013.301468 https://www.journals.uchicago.edu/doi/abs/10.1086/686466?mobileUi=0&journalCode=ssr
https://www.tandfonline.com/doi/abs/10.1080/10511482.2022.2077801
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Department of Social Services
Year: 2023 • Agency: Department of Social Services • Department: NA
Year
2023
Agency
Department of Social Services
Department
NA
Tool Name
SmartVAN / TargetSmart
Date First Use
2019/11
Updated
Yes
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The Mayor’s Public Engagement Unit (PEU) uses SmartVAN to manage outreach across a range of projects. SmartVAN provides functionality to create lists of potential clients to contact, collect personal information and survey responses from clients, and conduct outreach via phone banks and canvassing. SmartVAN also contains a frequently updated commercial dataset, provided by TargetSmart, of New York City residents and their demographic, contact, and other information. PEU uses this preloaded data to create outreach lists when data on existing clients or from partner agencies is unavailable or insufficient to meet the scope of the outreach project.
Purpose Desc
In 2023, PEU has the TargetSmart data within SmartVAN on a number of projects. PEU frequently uses the data to create lists of residents who live within certain zip codes that PEU wants to target for outreach. For example, PEU created lists using TargetSmart data to conduct door-knocking and phone banking outreach to New Yorkers identified as potentially eligible for the DOF Rent Freeze program based on TargetSmart data. In cases like these, TargetSmart’s determination of who lives in which zip codes as well as estimated income affects whether New Yorkers receive PEU outreach. Additionally, the algorithm that TargetSmart uses to match phone numbers to individuals impacts the type of outreach that New Yorkers receive.
Updated Desc
TargetSmart data available in SmartVAN is updated on a regular basis by the vendors.
Identifying Info
true
Data Training
Training data is part of vendor’s proprietary processes.
Data Input
Input data is part of vendor’s proprietary processes.
Data Output
The algorithmically-derived data that PEU accesses is the output of proprietary algorithmic processes developed and operated by TargetSmart. These algorithmic processes include matching multiple input datasets to determine residency, contact information, and demographics on New York City residents. SmartVAN also includes a number of algorithmically-determined likelihood scores, including scores for the likelihood that a household contains children under 18, etc.
Vendor Name
EveryAction and TargetSmart
Vendor Type
NA
Vendor Desc
EveryAction and TargetSmart jointly provide the SmartVAN product. EveryAction is the software provider. TargetSmart is the data provider. TargetSmart is the entity who applies algorithmic techniques. EveryAction provides access to this data through their platform.
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling; Matching
Department of Social Services
Year: 2023 • Agency: Department of Social Services • Department: NA
Year
2023
Agency
Department of Social Services
Department
NA
Tool Name
Splink
Date First Use
2023/03
Updated
Created in CY2023
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Splink is Python package used for entity resolution (i.e., deduplication) of records in which there is no unique identifier. It helps uses implement probabilistic matching.
Purpose Desc
In 2023, PEU moved the Tenant Helpline onto a live caller system using the Virtual Call Center and Salesforce. Prior to migrating data from the original system (which involved clients leaving voicemails), PEU needed to identify call records from the same client. PEU used the Splink package to identify records from the same caller even in cases with relatively sparse information and/or inconsistent data entry (e.g., different spellings of the same name).
Updated Desc
NA
Identifying Info
true
Data Training
PEU fine-tuned our use of the algorithm using a sample of the original Tenant Helpline data. There are a number of user-defined parameters that affect the results of the matching. PEU tested which parameters were most appropriate for our use case of very sparse data and conducted human quality control over its performance.
Data Input
The input data was information on individual Tenant Helpline callers. This included their names, phone numbers, address information, etc.
Data Output
The output of PEU’s usage of Splink was a set of callers identified as unique linked to one or more calls to the Tenant Helpline. This information was transformed and migrated to the new Salesforce system.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Matching
Fire Department
Year: 2023 • Agency: Fire Department • Department: NA
Year
2023
Agency
Fire Department
Department
NA
Tool Name
EMD Schedule Optimization Tool
Date First Use
2021/06
Updated
Yes
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Other
Population Type Individual
NA
Population Type Other
FDNY radio and assignment dispatcher employees
Website
NA
Tool Desc
The purpose of the tool is to provide Emergency Medical Dispatchers (EMD) staff a tool to optimally allocate call takers during a 24-hour period. The tool uses an expected number of incoming calls and the number of personnel scheduled to work in order to allocate the call takers to different shifts such that the supply of call takers exceeds the demand for call takers.
Purpose Desc
The algorithm requires two datasets. First, the tool requires the average number of medical calls per hour for a 24-hour period. Second, the tool requires a user to specify the number of call takers assigned to each tour. Based on these two inputs, the tool provides a projection of supply (call takers) versus demand (medical calls). Additionally, the tool can take the total number of available staff and optimally allocate them across tours to maximize the minimum difference between supply and demand. Based on these outputs, EMD officers can identify times during the day when call taker utilization is high and reallocate staff to accommodate.
Updated Desc
Front-end user interface added
Identifying Info
false
Data Training
This is an optimization model and was not “trained” using training data. The algorithm relies on actual historical data to determine average hourly medical calls.
Data Input
The tool requires an hourly count of medical calls arriving during a 24-hour period. Additional “data” requirements are input from the user depending on user-driven scenarios. For example, a user could specify five eight-hour tours per day (at different start times) rather than existing four tours (two eight-hour tours and two 12-hour tours).
Data Output
The algorithm outputs a projection of supply (call takers) versus demand (medical calls). Additionally, the tool can take the total number of available staff and optimally allocate them across tours to maximize the minimum difference between supply and demand.
Vendor Name
None
Vendor Type
NA
Vendor Desc
The tool was developed internally at FDNY in partnership with Columbia University's Industrial Engineering and Operations Research Department.
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Fire Department
Year: 2023 • Agency: Fire Department • Department: NA
Year
2023
Agency
Fire Department
Department
NA
Tool Name
EMS Hospital Suggestion Algorithm
Date First Use
2007/03
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Geographic space
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The EMS Hospital Suggestion Algorithm is used to determine the closest, appropriate hospital to the incident location based on the needs of a patient requiring transport.
Purpose Desc
The algorithm computes a list of hospitals in order of closest to furthest in time for each medical condition category as currently established. (For example, there is a list of hospitals computed in order of closest in time for all hospitals that accept General Emergency Department patients and for all hospitals that accept special conditions, such as burns). Depending on the medical needs category of the patient, the algorithm produces a pre-determined list of hospitals which is based on the location of the patient and then made available to the crew as a list of “closest, most appropriate hospitals.”
Updated Desc
NA
Identifying Info
false
Data Training
The EMS Hospital Suggestion algorithm relies on telematics data from the Department of Citywide Administrative Services city-owned vehicles collected between 2015 and mid-2016 to calibrate a network analysis model that derives incident to hospital transport times. The order of suggested hospitals are then compared with five years of historical EMS hospital transport data from before the COVID-19 pandemic (2015-2019) to validate and correct the network model.
Data Input
The inputs for the algorithm include the location and medical condition of the patient.
Data Output
The algorithm outputs the closest, most appropriate hospitals.
Vendor Name
None
Vendor Type
NA
Vendor Desc
This algorithm and the resulting output file that is used in our EMS CAD system to suggest hospitals was provided by Deccan International, until September 2020. The Department currently creates this file using a new algorithm, developed in-house by the Geographic Information Science (GIS) unit in conjunction with engineers from Columbia University.
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Fire Department
Year: 2023 • Agency: Fire Department • Department: NA
Year
2023
Agency
Fire Department
Department
NA
Tool Name
EMS Unit Suggestion Algorithm
Date First Use
2007/03
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Geographic space
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The Emergency Medical Services (EMS) Unit Suggestion Algorithm is used to determine which order of geographic regions (known as atoms) to search in order for the EMSCAD system to select an appropriate EMS unit for dispatch to an incident.
Purpose Desc
The algorithm computes a list of geographic regions (known as atoms) in order of closest to furthest in travel time for each atom in the city. This list of ordered atoms is the output of an algorithm that relies on a calibrated network model to derive travel time estimates. The output is an excel file which is converted into an EMSCAD-compatible file and loaded into the system for real-time unit selection capabilities. The file is generated and implemented as a 24/7 source file, meaning, the recommended search order is not currently varying by time of day. The Department is intending to implement time-of day search orders in the near future.
Updated Desc
NA
Identifying Info
false
Data Training
The EMS Unit Suggestion algorithm relies on historical FDNY CAD trip time data which is used to calibrate a network analysis model which derives atom-to-atom transport times.
Data Input
The input for the algorithm is a geographic location.
Data Output
The algorithm outputs a recommended EMS unit for dispatch.
Vendor Name
Deccan International
Vendor Type
NA
Vendor Desc
This algorithm and the resulting output file that is used in our EMS CAD system to suggest atom order for unit search is currently provided by a vendor, Deccan International.
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Fire Department
Year: 2023 • Agency: Fire Department • Department: NA
Year
2023
Agency
Fire Department
Department
NA
Tool Name
RBIS (Risk Based Inspection Program); ALARM (A Learning Approach to Risk Modeling)
Date First Use
2019/11
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Property; Other
Population Type Individual
NA
Population Type Other
civilian fatalities
Website
NA
Tool Desc
A Learning Approach to Risk Modeling (ALARM) creates risk scores for each building in the city. These scores are used to schedule our Fire Operations building inspections within the inspectable population of buildings in the city (~330,000 Building Identification Numbers (BINs)), as a part of the Risk-Based Inspection Program (RBIS).
Purpose Desc
ALARM is a combined approach using machine learning and risk ratios to assess the risk of a building for structural fire ignition (probability) and civilian fire injury/death (impact). The machine learning algorithm takes incident data, housing characteristics, and 311 data and creates a probability of structural fire ignition. This is combined with a civilian injury or death risk ratio for the building which is based on building characteristics, incident data and nearby felony crimes to create a risk score (range is one to nine), with one being highest risk and nine being lowest risk. Buildings are prioritized within each of the nine risk scores according to the residential population in each building.
Updated Desc
NA
Identifying Info
false
Data Training
In order to create the models, the team utilized a five-year incident dataset and reserved 99 percent of the data to train the probability model and 80 percent of the data to train the impact model.
Data Input
The ALARM risk score utilizes data from our fire and Emergency Medical Services (EMS) dispatch system, building characteristic data, 311 calls, felony crimes, census data and civilian injury data.
Data Output
The tool outputs a risk score from one (highest risk) to nine (lowest risk).
Vendor Name
None
Vendor Type
NA
Vendor Desc
ALARM was built in-house by a team of analysts from the Management, Analysis and Planning Bureau.
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Health + Hospitals
Year: 2023 • Agency: Health + Hospitals • Department: NA
Year
2023
Agency
Health + Hospitals
Department
NA
Tool Name
Adult Risk of IP/ED Utilization Score
Date First Use
2019/01
Updated
Yes
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The Adult Risk of Inpatient (IP)/Emergency Department (ED) Utilization Score predicts the number of days in the ED or IP setting that a patient may have in the coming year.
Purpose Desc
The Adult Risk of IP/ED Utilization Score predicts the number of days in the ED or IP setting that a patient may have in the coming year. It uses internal electronic medical record data covering past utilization, diagnoses, and documented behavioral health risk factors. Patients in the top 5 percent may be eligible to participate in the health home program, even if they are not Medicaid eligible.
Updated Desc
Updated our homelessness variable to incorporate additional sources of documentation within our system, and validated the algorithm against fiscal year 2021 data to evaluate continued fit.
Identifying Info
true
Data Training
This model was trained on NYC Health + Hospital population to create a utilization prediction method that was responsive to the need of safety-net patients. Most national algorithms are tuned on large claims dataset and may not be generalizable to uninsured patients. The model was developed with 70+ predictor variable and tuned using a LASSO continuous model. We assessed 70+ predictor variables on electronic health record data from CY Q3 2016-Q2 2017 to train the model (n=833,969). Data from CY Q3 2017-Q2 2018 was used to validate. This study was approved by NYC Health + Hospital Institutional Review Board (IRB) partner, BRANY IRB.
Data Input
The final model had 17 predictors. The top binary predictors for the final model were psychosis diagnosis (β=1.17), history of incarceration (β=0.47), antipsychotic medication prescription (β=0.40), and substance use disorder diagnosis (β=0.38). Top continuous predictors were inpatient visits (β=0.36), ED visits (β=0.34), and number of chronic conditions (β=0.21).
Data Output
Outputs of the model (top 5 percent flag) allow clinicians to connect patients with social work support or referrals to community organizations.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Mayor's Office
Year: 2023 • Agency: Mayor's Office • Department: NA
Year
2023
Agency
Mayor's Office
Department
NA
Tool Name
ElevenLabs Speech Synthesis - MO - COS
Date First Use
2023/06
Updated
Created in CY2023
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
ElevenLabs creates the most realistic, versatile and contextually-aware AI audio, providing the ability to generate speech in hundreds of new and existing voices in over 20 languages.
Purpose Desc
The tool was used to generate audio recordings of mayor�s voice delivering hiring hall/RiseUp concert messages in various languages to be used for hiring hall/RiseUp concert robo-calls.
Updated Desc
NA
Identifying Info
true
Data Training
Submitted audio recordings of mayor�s speech, as well as a live language sample, into the program.
Data Input
Script translated in desired languages (Spanish, Yiddish, Haitian Creole).
Data Output
Audio recording (MP3) of mayor�s voice speaking the script in the desired language (Spanish, Yiddish, Haitian Creole).
Vendor Name
ElevenLabs
Vendor Type
NA
Vendor Desc
Utilized two existing features offered by ElevenLabs (VoiceLab and Speech Synthesis).
Data 2022
NA
Vendor
NA
Analysis Type
Speech and language processing
Mayor's Office
Year: 2023 • Agency: Mayor's Office • Department: NA
Year
2023
Agency
Mayor's Office
Department
NA
Tool Name
Methodology for Poll Site Language Assistance - MO - CEC
Date First Use
2020/11
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Since no dataset is currently available that reliably captures the number of limited English proficient (LEP) registered voters for all program languages, the Civic Engagement Commission (CEC) uses the percentage of LEP citizens of voting age (CVALEP) as a substitute or proxy measure of need. CEC ranks the program-eligible languages in order of magnitude of CVALEP and distributes poll sites to each language based on its ranking (excluding CVALEP persons that speak languages served by the NYC Board of Elections (NYCBOE) in certain New York City counties). The number of poll sites that will receive services in any given language will depend on each language’s share of the total CVALEP in the population eligible to be served. For example, according to U.S. Census data, approximately 207,926 New Yorkers are CVALEP and speak a language that is served by this program. This proportionality approach allows CEC to balance goals of including diverse language communities as well as fair access to the total number of eligible voters within each language community. The program provides interpreters in program-eligible languages at poll sites based on U.S. Census data showing concentrations of CVALEP individuals who speak these languages and reside around each poll site. For each language, poll sites are chosen in descending order of concentration of CVALEP, until the language’s share is met. This process is repeated for each language, thereby including the poll sites with the highest concentration of CVALEP for each program-eligible language until that language’s share is met, and the total number of poll sites for which resources are allocated is reached. It may be possible, based on analysis of data, to reassign poll sites to languages with greater need; however, each language will receive a minimum of at least one poll site. Models used included the Thiessen polygon method to create a Voronoi diagram to determine CVALEP estimates.
Purpose Desc
This is a methodology for determining how the New York City Civic Engagement Commission (NYCCEC) will provide interpretation services at poll sites for limited English proficient voters. The methodology explains how the NYCCEC will identify the languages and locations in which interpretation services will be offered during the November 2020 election and beyond. These services supplement the interpretation assistance provided by NYC Board of Elections in several languages. Under the Charter, the NYCCEC can only provide interpretation services in a language if: (1) it is a designated citywide language; or (2) it is spoken by a greater number of LEP New Yorkers than the lowest ranked designated citywide language and at least one poll site has a significant concentration of speakers of such language with LEP. This methodology ensures service for all languages that are eligible under the Charter.
Updated Desc
NA
Identifying Info
false
Data Training
N/A
Data Input
For citywide estimates, this methodology uses current data from the American Community Survey (ACS) 2016-2020 five-year estimates. This methodology also uses the American Community Survey Census Tract 2016-2020 five-year Public Use Microdata Samples for poll site level analysis; this is the most current and accurate data available on resident New Yorkers at the neighborhood level. In addition, the methodology uses data from the Board of Elections on the location of election districts and poll sites.
Data Output
The algorithm estimates the number of citizens of voting age with Limited English Proficiency for each program-eligible language who could report to each polling site.
Vendor Name
None
Vendor Type
NA
Vendor Desc
The tool was designed with the support of the Mayor’s Office of Data Analytics which is currently part of the Office of Technology and Innovation.
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
Mayor's Office
Year: 2023 • Agency: Mayor's Office • Department: NA
Year
2023
Agency
Mayor's Office
Department
NA
Tool Name
Scorecard Blockface Sampling Algorithm - MO - Operations
Date First Use
2022/03
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Geographic space
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The Scorecard program sends inspectors across New York City to rate street and sidewalk cleanliness. The sampling algorithm creates a monthly list of blocks for inspectors to visit and rate.
Purpose Desc
The primary goal of the algorithm is to produce a sample of blockfaces that is statistically sound and geographically representative. This list is used to rate street and sidewalk cleanliness citywide, as well as by borough and DSNY district.
Updated Desc
NA
Identifying Info
false
Data Training
N/A
Data Input
The blockface sample is selected from the Pavement Edge File, which is part of the New York City Planimetric Database managed by the Office of Technology and Innovation. Sampling is weighted towards blockfaces in high-density areas and includes extra sampling of blockfaces in Business Improvement Districts (BIDs). It also takes into account the linear miles of street within a DSNY District.
Data Output
A count of blockfaces that are statistically representative of our target areas throughout the city.
Vendor Name
Legacy Mayor’s Office of the Chief Technology Officer
Vendor Type
NA
Vendor Desc
The sampling algorithm was developed by the former Mayor’s Office of the Chief Technology Officer in partnership with the Mayor’s Office of Operations.
Data 2022
NA
Vendor
NA
Analysis Type
Optimization
New York Police Department
Year: 2023 • Agency: New York Police Department • Department: NA
Year
2023
Agency
New York Police Department
Department
NA
Tool Name
Facial Recognition Technology
Date First Use
2011/10
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Tool which may help investigators identify unknown subjects in law enforcement investigations.
Purpose Desc
Facial recognition is a digital technology that NYPD uses to compare images obtained during investigations with lawfully possessed arrest and parole photos. The tool analyzes an uploaded image, known as a probe image, and searches and compares against the image repository. The purpose of the tool is to enhance law enforcement’s ability to investigate criminal activity as well as identify deceased persons and missing persons. When used in combination with human analysis and additional investigation, facial recognition technology is a valuable tool in solving crimes and increasing public safety.
Updated Desc
NA
Identifying Info
true
Data Training
Training data is proprietary to the vendor.
Data Input
If NYPD investigators obtain a still image depicting a face of an unknown individual during an investigation, the image can be submitted for facial recognition analysis in accordance with NYPD facial recognition policy. Known as a probe image, NYPD facial recognition software compares the image to a controlled and limited group of lawfully obtained photos called the photo repository.
Data Output
The facial recognition software will generate a pool of possible match candidates for review by trained Facial Identification Section investigators.
Vendor Name
Dataworks
Vendor Type
NA
Vendor Desc
Software developed and maintained by vendor
Data 2022
NA
Vendor
NA
Analysis Type
Computer vision; Matching
New York Police Department
Year: 2023 • Agency: New York Police Department • Department: NA
Year
2023
Agency
New York Police Department
Department
NA
Tool Name
Patternizr
Date First Use
2016/12
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals; Property; Geographic space; Other
Population Type Individual
NA
Population Type Other
Crime Classification
Website
NA
Tool Desc
Aids crime analysis in detection of potential crime patterns.
Purpose Desc
Patternizr compares features of crimes and finds ones that are similar and may be part of a crime pattern. Analysts will look at the candidate crimes and suggest the formation of crime patterns to a pattern identification module. If a pattern is formed, detectives often consolidate the investigative efforts (e.g., one detective investigates all the crimes in the pattern.) The report filters non-normal trends into a spreadsheet and displays year-over-year counts of crimes that have non-normal trends. The tool requires a human user to evaluate the output data to see if complaints identified as similar are, in fact, connected to a pattern.
Updated Desc
NA
Identifying Info
false
Data Training
Separate models were trained for each of three different crime types (burglaries, robberies, and grand larcenies). These crime types have a sufficient corpus of prior manually identified patterns for use as training examples. This corpus consists of approximately 10,000 patterns between 2006 and 2015 from each crime type. A portion of this corpus includes complaint records where the same individual was arrested for multiple crimes of the same type within a span of two days.
Data Input
The input data is a candidate crime and its features. A complaint describes details of the crime, including the date and time (which can be a range if the precise time of occurrence is unknown), location, crime subcategory, modus operandi, and suspect information. This information is used to calculate the five types of crime-to-crime similarities used as features by Patternizr: location, date-time, categorical, suspect and unstructured text.
Data Output
Probability that a complaint is connected to a pattern.
Vendor Name
None
Vendor Type
NA
Vendor Desc
The tool was developed by data scientists and analysts at NYPD. Contractors and NYPD personnel integrated it into the Domain Awareness System. Personnel in Crime Control Strategies work with the Information Technology Bureau to maintain the tool.
Data 2022
NA
Vendor
NA
Analysis Type
Matching
New York Police Department
Year: 2023 • Agency: New York Police Department • Department: NA
Year
2023
Agency
New York Police Department
Department
NA
Tool Name
ShotSpotter
Date First Use
2015/03
Updated
Yes
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Geographic space
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Provides acoustic gunshot detection to assist with emergency call response
Purpose Desc
Provides acoustic gunshot detection to assist with emergency call response. The tool supports patrol operations in alerting units to potential gunfire and enhances investigations involving firearms.
Updated Desc
Routine Maintenance
Identifying Info
false
Data Training
Training data is proprietary to the vendor.
Data Input
Specialized software analyzes audio signals for potential gunshots.
Data Output
The tool determines the location of the sound source, and once classified as potential gunfire sends the incident to acoustic experts for additional analysis. Notifications are sent for confirmed gunfire. ShotSpotter activations may result in evidence collection that can enhance case investigations. Problematic locations identified through alerts may require additional resource deployment and/or investigations.
Vendor Name
Shotspotter
Vendor Type
NA
Vendor Desc
Software developed and maintained by vendor
Data 2022
NA
Vendor
NA
Analysis Type
Matching
NYC Public Schools
Year: 2023 • Agency: NYC Public Schools • Department: NA
Year
2023
Agency
NYC Public Schools
Department
NA
Tool Name
Eureka! Chatbot
Date First Use
2023/08
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The Azure Cognitive Services technology and chatbot (internally branded as "Eureka!" has been configured and deployed in August 2023 to be the first response to calls to the NYCPS IT Service Desk. It accesses scripts to handle four common reasons for a user to call or contact the service desk - Password Reset, Create a Ticket, Ticket Status, Request for Information. The chatbot accesses pre-defined scripts to respond to user voice or text input. The user's request is either serviced, completed and closed by the chatbot, or the user is given the option (at any time) to connect to a live agent.
Purpose Desc
The tool is used to respond to common IT service desk requests - Password Reset, Create a Ticket, Ticket Status, Request for Information. Users can access the tool by phone, by computer through the DOE Support Hub application, and from links from MS Teams and other DOE systems, such as TeachHub.
Updated Desc
NA
Identifying Info
true
Data Training
Pre-defined scripts designed to respond to four common requests to the IT Service Desk.
Data Input
A voice call or text-based chat session initiated by a user and responded to by the Eureka! chatbot before being handled by a human Service Desk agent.
Data Output
The chatbot generates responses to user-entered prompts based on the training data, or forwards the inquiry to a human Service Desk agent. Since its launch in August 2022, the chatbot handles an average of 1,500 calls and 300 web-based inquiries each day. Approximately 30 percent of the voice calls and 10 percent of the web-based queries have been handled completely by Eureka! without being forwarded to a human Service Desk agent.
Vendor Name
Nagarro and Microsoft
Vendor Type
NA
Vendor Desc
Developed by an IT services vendor (Nagarro) using Microsoft Cognitive services.
Data 2022
NA
Vendor
NA
Analysis Type
Speech and language processing
NYC Public Schools
Year: 2023 • Agency: NYC Public Schools • Department: NA
Year
2023
Agency
NYC Public Schools
Department
NA
Tool Name
MySchools
Date First Use
2018/08
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The tool utilizes the Gale-Shapley deferred acceptance algorithm to match applicants to schools. This algorithm has been in existence for many years, used internationally for various purposes. Perhaps most common is its use in the National Resident Matching Program for medical school students.

Deferred acceptance works as an iterative series of steps: students and programs are tentatively matched in each step, but nothing is finalized until the algorithm terminates (hence the deferred).
1. Each student “proposes” to their first choice
• Programs assign seats to students one at a time
• When all seats are filled, programs may reject previously accepted students in favor of new applications from students they prefer (e.g., students with a better lottery number)
• Remaining students are rejected
2. Students rejected in the last step “propose” to the next choice on their list
3. The algorithm terminates when all students are matched or have proposed to all the programs they listed
Purpose Desc
MySchools is an application used to house online school directories, collect application choices, and run the admissions matching algorithm that is used for all centralized admissions processes (3K, pre-K, Gifted & Talented, middle school, and high school). The tool encompasses a family-facing portal, a school-facing portal, and an administrative portal.
Updated Desc
NA
Identifying Info
false
Data Training
The algorithm was already widely recognized for its advantages prior to adoption in New York City. The DOE consulted with a team of researchers at MIT who had been closely involved in its initial creation when we adopted it.
Data Input
Student biographical information (e.g., home address, poverty status, home language), student academic information (e.g., course grades, state test scores), and student school records (e.g., sending school).
Data Output
The algorithm outputs a school match for each student.
Vendor Name
Blenderbox
Vendor Type
NA
Vendor Desc
We have a five year contract with the agency Blenderbox who designed the application and implemented the algorithmic matching functionality. The work is meant to transition to be run in-house, by the Division of Instructional and Information Technology (DIIT) within the Department of Education, by the end of the contract. The team at DIIT has already begun to takeover maintenance and development of the tool.
Data 2022
NA
Vendor
NA
Analysis Type
Matching
NYC Public Schools
Year: 2023 • Agency: NYC Public Schools • Department: NA
Year
2023
Agency
NYC Public Schools
Department
NA
Tool Name
NYCDOE APPR Measures of Student Learning (MOSL) Growth Model
Date First Use
2013/09
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The growth model uses a variety of student-level (assessment scores, English Language Learner, Disability, and Economic Disadvantage indicators), classroom-level (e.g. percent Students With Disabilities), and school-level data (e.g. percent English Language Learners, percent Students With Disability, average prior achievement, school type) to estimate/predict a student's score on one of many possible course-culminating assessments. These predicted scores are used to either 1) identify "peer groups" of students, from which student growth percentiles (SGPs) are determined, or 2) compared to actual scores to determine student credit values. These units (SGPs or credit values) are then weight-averaged to generate an educator-level result - the MOSL Rating. The MOSL Rating is combined with the MOTP Rating to produce an Overall Rating. Per state law 3012-d, annual ratings “shall be a significant factor in HR decisions.” This is often implemented by making ratings a qualifying/disqualifying element in decision-making concerning employment, tenure, salary, and other professional opportunities.
Purpose Desc
In accordance with New York state law and New York State Education Department (NYSED) regulations, the Department developed and maintains a "growth model" to produce Measures of Student Learning (MOSL) ratings for use in annual professional performance reviews (APPR) for teachers and principals. The MOSL ratings are combined with Measures of Teaching/Leadership Practice (MOTP/MOLP) ratings to produce an annual Overall Rating for each eligible educator.
Updated Desc
NA
Identifying Info
false
Data Training
The growth model process is employed in both retrospective and prospective ways. In the retrospective version, the results are determined entirely within-sample. In the prospective version, the coefficients of the model are estimated on multiple prior years of data.
Data Input
The growth model makes use of three types of data: (1) students’ end-of-year assessment scores, (2) enrollment and attendance records that link students to teachers and schools, and (3) historical academic and demographic information used to identify groups of similar students.
Data Output
The model outputs an estimate of a student's score on a course-culminating assessment.
Vendor Name
Education Analytics
Vendor Type
NA
Vendor Desc
Education Analytics provides technical assistance and quality assurance for the growth model.
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
NYC Public Schools
Year: 2023 • Agency: NYC Public Schools • Department: NA
Year
2023
Agency
NYC Public Schools
Department
NA
Tool Name
NYCDOE APPR Measures of Teaching/Leadership Practice (MOTP/MOLP) Calculation
Date First Use
2013/10
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Throughout a school year, evaluators observe teachers/principals multiple times and use a rubric to provide a numerical rating on one or more rubric components. These rubric component scores are then weight-averaged according to collectively bargained rules to produce an MOTP/MOLP Rating. The MOTP/MOLP Rating is combined with the MOSL Rating to produce an Overall Rating for each eligible educator. Per state law 3012-d, annual ratings “shall be a significant factor in HR decisions.” This is often implemented by making ratings a qualifying/disqualifying element in decision-making concerning employment, tenure, salary, and other professional opportunities.
Purpose Desc
In accordance with New York state law and New York State Education Department (NYSED) regulations, the Department developed and maintains databases and calculation rules to produce Measures of Teaching/Leadership Practice (MOTP/MOLP) ratings for use in annual professional performance reviews (APPR) for teachers and principals. The MOTP/MOLP ratings are combined with Measures ofStudent Learning (MOSL) ratings to produce an annual Overall Rating for each eligible educator.
Updated Desc
NA
Identifying Info
false
Data Training
Pilot data prior to program launch was used to inform the weights assigned to various rubric components. However, the weights are ultimately determined via collective bargaining.
Data Input
Rubric component numerical ratings.
Data Output
The model outputs a score for teachers and principals.
Vendor Name
None
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Predictive modeling
NYC Public Schools
Year: 2023 • Agency: NYC Public Schools • Department: NA
Year
2023
Agency
NYC Public Schools
Department
NA
Tool Name
Open Gen AI and Teaching Assistant Tool
Date First Use
2023/05
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The generative AI system using large language models was a system custom-built by DIIT using advanced Microsoft technologies to create a set of generative AI tools. To date, two tools have been built. One is named “Open Gen AI” - it accesses a large language model (currently GPT 3.5) to provide responses to a broad range of prompts. The other is named “Teaching Assistant for Algebra” - it accesses specific Algebra-focused content to provide responses to prompts related to Algebra.
Purpose Desc
The tool is used to generate responses to prompts entered by a student or teacher, requesting the generative AI tool to compose a text response to a text input.
Updated Desc
NA
Identifying Info
true
Data Training
For the Open Gen AI tool, the ChatGPT large language model is used as the "trained data". For the Teaching Assistant for Algebra tool, the LLM has been trained exclusively on curriculum from Illustrative Math.
Data Input
Prompts provided by the users of the system.
Data Output
The output data for the Open Gen AI tool is the response generated by the ChatGPT LLM. The output data for the Teaching Assistant is the response generated by specifically developed LLM using the Illustrative Math curriculum.
Vendor Name
Microsoft
Vendor Type
NA
Vendor Desc
Microsoft provided technical guidance for their emerging generative AI technology and built some small module of code for the specific NYCPS Gen AI and Teaching Assistant use cases.
Data 2022
NA
Vendor
NA
Analysis Type
Speech and language processing
Office of Chief Medical Examiner
Year: 2023 • Agency: Office of Chief Medical Examiner • Department: NA
Year
2023
Agency
Office of Chief Medical Examiner
Department
NA
Tool Name
STRMix
Date First Use
2017/01
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
STRmix™ combines sophisticated biological modeling and standard mathematical processes to interpret a wide range of complex DNA profiles. Using well-established statistical methods, the software builds millions of conceptual DNA profiles. It grades them against the evidential sample, finding the combinations that best explain the profile. A range of likelihood ratio options are provided for subsequent comparisons to reference profiles. Using a Markov Chain Monte Carlo engine, STRmix™ models any types of allelic and stutter peak heights as well as drop-in and drop-out behavior. It does this rapidly, accessing evidential information previously out of reach with traditional methods. STRmix™ is supported by comprehensive empirical studies with its mathematics readily accessible to DNA analysts, so results are easily explained in court.
Purpose Desc
STRMix is a probabilistic genotyping tool that is used to analyze mixtures of DNA profiles to help associate the crime scene evidence to potential victims or suspects of crimes.
Updated Desc
NA
Identifying Info
true
Data Training
Training data was not used in the sense of AI software. The OCME performed thousands of tests using the software to validate it for optimum use with our current laboratory standard operating procedures and genetic analyzers.
Data Input
Forensic DNA profiles from crime scenes as well as the DNA profiles from victims and suspects of crimes.
Data Output
The output is a deconvolution of genotype probability distribution that lists all of the accepted genotype sets and their associated weights. These weights can take any value from 0 to 1.
Vendor Name
NicheVision Forensics, LLC
Vendor Type
NA
Vendor Desc
The software has been developed by New Zealand Crown Institute of Environmental Science and Research (ESR) with Forensic Science South Australia. The developer assisted OCME in analyzing and interpreting our data during the validation of the software.
Data 2022
NA
Vendor
NA
Analysis Type
Other: STRmix is a forensic DNA analysis software program that uses a probabilistic genotyping algorithm to interpret complex DNA profiles, such as those from mixed samples that contain DNA from multiple contributors. Specifically, STRmix uses a continuous probabilistic modeling approach called Markov chain Monte Carlo (MCMC) analysis.
Office of Technology and Innovation
Year: 2023 • Agency: Office of Technology and Innovation • Department: NA
Year
2023
Agency
Office of Technology and Innovation
Department
NA
Tool Name
MyCity Chatbot
Date First Use
2023/09
Updated
Created in CY2023
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals; Group, organization, or business
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The NYC MyCity chatbot is a beta AI-powered chatbot that provides information and access to services for residents and businesses in New York City.
Purpose Desc
The NYC MyCity chatbot is a beta AI-powered chatbot that provides information and access to services for residents and businesses in New York City. It’s currently focused on two main areas: Business Services and MyCity Basics. The chatbot provides information on starting or operating a business in New York City, answers questions about permits, licenses, regulations, and other business requirements, and connects users with relevant resources and support services. It also offers information on various city services and benefits, and helps users find resources related to childcare, career, and other areas. The chatbot is using Microsoft’s Azure AI technology and OpenAI’s ChatGPT 3.5 Turbo LLM.
Updated Desc
NA
Identifying Info
false
Data Training
Training data is proprietary to the vendor.
Data Input
Text queries are input by the user on the MyCity portal.
Data Output
The tool produces text responses with references based on information from Business Services and MyCity Basics.
Vendor Name
Microsoft, Nuvalence
Vendor Type
NA
Vendor Desc
Microsoft provides Cloud-based ChatGPT services and Nuvalence was the professional services vendor for implementation.
Data 2022
NA
Vendor
NA
Analysis Type
Speech and language processing
School Construction Authority
Year: 2023 • Agency: School Construction Authority • Department: NA
Year
2023
Agency
School Construction Authority
Department
NA
Tool Name
GitHub Copilot
Date First Use
2023/05
Updated
No
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
Individuals
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
GitHub Copilot is an AI-powered code assistant that provides suggestions for whole lines or blocks of code in a wide range of programming languages. It leverages a vast codebase and machine learning to improve coding efficiency, helping programmers by autocompleting code snippets and offering context-appropriate code suggestions.
Purpose Desc
GitHub Copilot is primarily used by some of our software developers as an advanced coding assistant within our agency. Its role is to augment and streamline the coding process for our software development projects. By providing real-time code suggestions and completions, it reduces the time developers spend on routine coding tasks, allowing them to focus on more complex aspects of software development.

The tool functions by analyzing the context of the code being written and suggesting relevant, syntactically correct code snippets. This includes generating code for standard programming patterns, filling in boilerplate code, and offering solutions to simple programming queries. It's important to note that while GitHub Copilot assists in the coding process, final decisions on the code's implementation and its use in any software or application rest solely with our human developers. The tool's suggestions are always reviewed and potentially modified by our team to ensure they meet our specific requirements and standards. Therefore, GitHub Copilot acts as a support tool in the decision-making process of software development rather than a decisive entity.
Updated Desc
NA
Identifying Info
false
Data Training
GitHub Copilot was developed by OpenAI and trained using a large corpus of public source code available on GitHub. This training data includes a wide variety of code in multiple programming languages, along with associated comments and documentation. The data encompasses a broad range of coding styles, patterns, and solutions across different software development projects.
Data Input
When in use, GitHub Copilot analyzes the code that a developer is currently writing. This input data consists of the programming language syntax, structure, and any comments or context within the code file. The tool also takes into account the specific coding task, patterns, and functions that the developer is working on. This real-time data is essential for the tool to provide relevant and context-appropriate coding suggestions.
Data Output
The output data from GitHub Copilot includes suggested lines or blocks of code that align with the input data provided by the developer. These suggestions are generated based on the patterns, structures, and coding practices learned from the training data. The output is designed to seamlessly integrate with the existing code, offering syntactically correct and contextually relevant code completions.
Vendor Name
GitHub, a subsidiary of Microsoft
Vendor Type
NA
Vendor Desc
NA
Data 2022
NA
Vendor
NA
Analysis Type
Speech and language processing
Administration for Children's Services
Year: 2022 • Agency: Administration for Children's Services • Department: NA
Year
2022
Agency
Administration for Children's Services
Department
NA
Tool Name
Repeat Maltreatment (RM) model
Date First Use
2017/07
Updated
NA
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
NA
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Predictions of Repeat Maltreatment (identifying the likelihood of being involved in a future indicated investigation within the next 24 months) are based on Machine Learning methodology and are calculated for all children receiving prevention services from ACS prevention service providers. The prediction is made with the assumption that the case is closed the day the model is run. A prevention case is assigned a numeric likelihood of an indicated investigation based on a New York State Central Register (SCR) within 24 months from the end of a prevention service. This model was formerly known as the Service Terminaton Conference (STC) model and was used by Preventive Services managers to identify whether or not a case should have ACS or provider-agency facilitation at service termination conference.
Purpose Desc
The RM model was initially used to prioritize ACS facilitation of prevention termination conferences, so that ACS could be certain that services had been provided in these cases. When a family is ready to exit ACS prevention services, an end of services conference is required (known as a "Service Termination Conference"). However, Service Termination Conferences are generally no longer facilitated by ACS and this prioritization is no longer necessary.

The RM model is also used to group prevention providers into quartiles to assess their performance in comparison to others in the same quartile, based on the service-need/risk levels of the families they've served during the previous year. This is a retrospective analysis for performance management.
Updated Desc
NA
Identifying Info
NA
Data Training
NA
Data Input
NA
Data Output
NA
Vendor Name
NA
Vendor Type
NA
Vendor Desc
NA
Data 2022
Training data: ACS trained the model on ACS historic administrative data about closed investigations from July 2009 to December 2015. Training set included about 136,982 observations. The model was tested on closed investigations from January 2016 to December 2017 with 48,771 observations.

Input data: Predictions are based on administrative data about prior and current child welfare involvement including SCR investigations and time spent in foster care. Only ACS administrative data are used in the model.
Vendor
None
Analysis Type
NA
Administration for Children's Services
Year: 2022 • Agency: Administration for Children's Services • Department: NA
Year
2022
Agency
Administration for Children's Services
Department
NA
Tool Name
Severe Harm Predictive model
Date First Use
2018/05
Updated
NA
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
NA
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Predictions of Severe Harm (identifying likelihood of substantiated allegations of physical or sex abuse within the next 24 months) are based on Machine Learning methodology and are calculated for all children involved in active investigations early in the investigation (day 10). An investigation is assigned a numeric likelihood of this outcome based on the child in the case with the highest likelihood. The ACS Quality Assurance unit in the Division of Child Protection reviews about 3,000 active investigations annually, selecting those with the highest likelihood of severe harm. If the Quality Assurance review team identifies gaps in routine, required documentation or practice, the team speaks with the field office conducting the investigation and follows up to make certain these gaps have been addressed. No staff in Quality Assurance unit or in the investigative unit sees the scores. The model only supports the decision about which investigations are prioritized for review by the Quality Assurance unit.
Purpose Desc
The Quality Assurance Unit in the Division of Child Protection at ACS has the capacity to review about 3,000 investigation cases out of about 50,000 investigations annually. ACS developed this predictive model to support the selection of cases for Quality Assurance review. Open investigations involving children with the greatest likelihood to experience future severe harm -- substantiated allegations of physical or sex abuse in the following 24 months -- are selected for review. The tool does not support decisions about services or interventions for individuals or families involved with ACS, beyond the selection of the case for this additional Quality Assurance review.
Updated Desc
NA
Identifying Info
NA
Data Training
NA
Data Input
NA
Data Output
NA
Vendor Name
NA
Vendor Type
NA
Vendor Desc
NA
Data 2022
Training data: ACS trained the model on ACS historic administrative data about closed investigations from January 2013 to March 2014. Training set included about 119,668 observations. The model was tested on closed investigations from April 2014 to December 2014 with 82,331 observations.

Input data: Predictions are based on administrative data about prior and current child welfare involvement including investigations triggered by a New York State Central Register (SCR) call and time spent in foster care. Only ACS administrative data are used in the model.
Vendor
None
Analysis Type
NA
Department of Consumer and Worker Protection
Year: 2022 • Agency: Department of Consumer and Worker Protection • Department: NA
Year
2022
Agency
Department of Consumer and Worker Protection
Department
NA
Tool Name
Route Automation
Date First Use
2020/07
Updated
NA
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
NA
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Inspection Supervisor selects an inspector, enters a date and the number of businesses to be inspected, and the geographic area to be considered. The system identifies businesses in the selected area and assigns them to the route based on inspection priority until the number of businesses entered has been reached. Then the tool runs a Simulated Annealing Algorithm to optimize the order businesses appear on the route based on proximity and method of travel.
Purpose Desc
DCWP inspectors conduct inspections based on a route, or list of businesses to be inspected on a specific day, which must be pre-approved by their supervisor. The Route Automation tool generates a route for an inspector on a specific date based on configuration variables and geographic area. All routes generated by the tool still require supervisor review and approval.
Updated Desc
NA
Identifying Info
NA
Data Training
NA
Data Input
NA
Data Output
NA
Vendor Name
NA
Vendor Type
NA
Vendor Desc
NA
Data 2022
Training data: The tool is not trained in an AI or Machine Learning sense.

The tool makes decisions based on configuration tables, businesses and their licenses and inspection and violation history, and uses a Simulated Annealing Algorithm to optimize the order in which businesses appear on a route.

Input data: Inspection date, business category and address, licenses held (if any), last inspection date and type, violation history (if any), date of inspection request or license application/renewal (if applicable), Inspection Unit (of the inspector), and geographic area.
Vendor
The tool was designed and built by PruTech, an outside vendor contracted to design and build DCWP's Automated Inspection Management System (AIMS) and its accompanying Mobile Enforcement platform. The tool is part of the AIMS system.
Analysis Type
NA
Department of Correction
Year: 2022 • Agency: Department of Correction • Department: NA
Year
2022
Agency
Department of Correction
Department
NA
Tool Name
Housing Unit Balancer (HUB)
Date First Use
2017/04
Updated
NA
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
NA
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The HUB is comprised of two functions: (1) a classification tool based on decision trees that determines an individual's propensity for violence, and (2) a housing area risk assessment, which utilizes advanced predictive analytics (i.e., neural networks) to determine optimal housing areas based on the classification scores of people in custody. The primary operational use of the HUB is for the classification score, which is used to track populations and optimize housing arrangements. The last day the department utilized the HUB system for classification purposes was January 31, 2022.
Purpose Desc
The Housing Unit Balancer (HUB) is used for informing housing decisions made by operational staff designed to produce less conflict in housing areas.
Updated Desc
NA
Identifying Info
NA
Data Training
NA
Data Input
NA
Data Output
NA
Vendor Name
NA
Vendor Type
NA
Vendor Desc
NA
Data 2022
Training data:
Classification function: The training set comprised of over 60,000 records of people in custody and was retrieved in 2014. The test set was comprised of over 4,000 people in custody. These sets were used to create and validate the ultimate solution for the classification scoring of people in custody. Housing Area Risk Assessment function: The training set was incident data for 24 months (01/13-01/15) and the test set was incident data for 6 months (01/15-06/15).

Input data:
Classification function: (1) past institutional conduct, (2) top criminal charge, (3) security risk group affiliation, (4) age, (5) Brad H indicator. Housing Area Risk Assessment function: composition levels of (1) people in custody, (2) security risk group members, (3) security risk groups, (4) housing types, (5) age, (6) length of stay, and (7) duration of stay in housing areas.
Vendor
The tool was created by McKinsey & Company and implemented in 2017
Analysis Type
NA
Department of Health and Mental Hygiene
Year: 2022 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2022
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
BioNumerics Division of Disease Control
Date First Use
2017/09
Updated
NA
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
NA
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
BioNumerics is software used to store and analyze sequencing data from bacterial pathogens that are implicated in food outbreaks. In short, bacterial isolates derived from clinical and environmental sources are received by the NYC Public Health Laboratory where they are processed, tested, and ultimately sequenced to identify clusters of disease.
Purpose Desc
BioNumerics is used to 1) re-assemble the bacterial genome (since the sequencing process involves fragmenting the bacterial DNA and then amplifying it into millions of pieces) 2) identify the genus, species, and serotype of the bacterial isolate 3) perform quality control checks to ensure the sequence meets certain quality standards 4) perform whole genome multi-locus sequence typing (wgMLST, a technique used to type bacteria based on their genetic code) 5) perform cluster analysis for cases related to one another based upon case definitions recommended by the CDC. This information is then communicated to partners including foodborne epidemiologists at the Bureau of Communicable Disease, who investigate all reported cases of foodborne disease, with those investigations potentially resulting in restaurant inspections, closures, and food recalls.
Updated Desc
NA
Identifying Info
NA
Data Training
NA
Data Input
NA
Data Output
NA
Vendor Name
NA
Vendor Type
NA
Vendor Desc
NA
Data 2022
Nucleic acid sequences recovered from pathogen genomes using high throughput sequencing.
Vendor
None
Analysis Type
NA
Department of Health and Mental Hygiene
Year: 2022 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2022
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
ChoiceMaker (CM) Division of Disease Control
Date First Use
2003/06
Updated
NA
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
NA
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
CM is a record-matching tool that identifies duplicate records belonging to the same individual.
Purpose Desc
CM is used by BOI and Healthy Homes to identify duplicate immunization and lead records. The outputs produced by CM are used in ongoing manual and automated deduplication processes (record merging).
Updated Desc
NA
Identifying Info
NA
Data Training
NA
Data Input
NA
Data Output
NA
Vendor Name
NA
Vendor Type
NA
Vendor Desc
NA
Data 2022
CM uses demographic data (e.g., names, DOB, address, identifiers) and health event data (e.g., date and type of event) from BOI's Citywide Immunization Registry (CIR) and Healthy Homes' LeadQuest registry in its evaluation. The program outputs a series of record pairs and a match probability for each pair.
Vendor
A vendor was involved in the development of the program initially. CM is now available as an open-source program. The DOHMH implementation is maintained by HLN Consulting.
Analysis Type
NA
Department of Health and Mental Hygiene
Year: 2022 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2022
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
GATK Division of Disease Control
Date First Use
2017/10
Updated
NA
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
NA
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
A suite of tools for variant calling and filtering after sequence alignment. It uses naive Bayesian to qualify aligned bases as sequence or erroneous data, which would be excluded from the final genomic sequence.
Purpose Desc
Used to identify mutations and other differences in sequences in microbial genomes. Can be used to determine different characteristics of microbial genomes.
Updated Desc
NA
Identifying Info
NA
Data Training
NA
Data Input
NA
Data Output
NA
Vendor Name
NA
Vendor Type
NA
Vendor Desc
NA
Data 2022
Assembled nucleic acid sequences recovered from pathogen genomes using high throughput sequencing, and quantitative data that represent the quality of the values
Vendor
None
Analysis Type
NA
Department of Health and Mental Hygiene
Year: 2022 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2022
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
Improving Foodborne Disease Outbreak Detection by Incorporating Complaints Identified in Social Media Data
Date First Use
2016/11
Updated
NA
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
NA
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Foodborne disease outbreaks are identified through many mechanisms. Restaurant associated outbreaks are often identified through complaints received via NYC’s 311 non-emergent information system, however not all individuals report to 311. The New York City Department of Health and Mental Hygiene (NYC DOHMH) in collaboration with Columbia University developed a text classifier program which monitors Yelp and Twitter data to identify complaints of foodborne illness, which was supported by grants from the Alfred P Sloan Foundation and the National Science Foundation. These data are used in addition to complaint data received through NYC’s 311 system to identify and respond to foodborne disease outbreaks.
Purpose Desc
The model uses data from Yelp restaurant reviews and Twitter data that is available on Twitter's publicly available API. The classifiers assign a "sick score" to each Yelp review or tweet indicating the likelihood that the review or tweet pertains to foodborne illness. The sick score is based on whether the review/tweet contains key words indicative of foodborne illness ("e.g. vomit"); the Yelp classifier also incorporates if the review indicates that multiple people became sick and if the review indicates a time between eating at a restaurant and illness onset (incubation period) that is consistent with foodborne illness. Each review and tweet with a sick score greater than or equal to a threshold value are reviewed and annotated by DOHMH foodborne disease epidemiology and environmental health staff to determine if the review/tweet was actually reporting foodborne illness possibly associated with a NYC restaurant; if yes, Yelp messages are sent to Yelp reviewers, requesting that they contact DOHMH, and a Twitter message with a survey link is tweeted back to Twitter users to confirm foodborne illness. Data from annotations are used to improve classifier performance. Foodborne disease complaints identified through Yelp and Twitter are combined with foodborne disease complaints reported to 311 to improve efficiency of outbreak detection.
Updated Desc
NA
Identifying Info
NA
Data Training
NA
Data Input
NA
Data Output
NA
Vendor Name
NA
Vendor Type
NA
Vendor Desc
NA
Data 2022
Training data was used in the development of both the Yelp and Twitter classifiers. Columbia University provided DOHMH with sample datasets for Yelp reviews and tweets, which were reviewed and annotated by DOHMH staff. For Yelp, training data was focused on the following: 1) if the review indicated foodborne illness, 2) if the incident occurred in the past 30 days, 3) if multiple people were sick and 4) if the incubation period was consistent with foodborne illness. For Twitter, training data was focused on if the tweet was indicating foodborne illness and if the incident occurred in NYC.
Vendor
DOHMH staff, including Bureau of Communicable Disease, Office of Environmental Investigations, and Division of Informatics and Information Technology & Telecommunications and Columbia University are involved in making decisions about the tool. Columbia University Department of Computer Science professors and doctoral students maintain the classifier. The project was previously funded by the Alfred P Sloan Grant, for which The Fund for Public Health in New York provided administrative support and grant management to DOHMH. This support and management ended at the completion of the grant in 2021.
Analysis Type
NA
Department of Health and Mental Hygiene
Year: 2022 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2022
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
IQTREE Division of Disease Control
Date First Use
2020/05
Updated
NA
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
NA
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
IQTREE uses maximum-likelihood regression to create phylogenetic trees from genomes.
Purpose Desc
Produces phylogenetic trees which are used to rule in or out SAR-CoV-2 sequences in outbreaks.
Updated Desc
NA
Identifying Info
NA
Data Training
NA
Data Input
NA
Data Output
NA
Vendor Name
NA
Vendor Type
NA
Vendor Desc
NA
Data 2022
Assembled nucleic acid sequences recovered from pathogen genomes using high throughput sequencing as input data to this tool
Vendor
None
Analysis Type
NA
Department of Health and Mental Hygiene
Year: 2022 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2022
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
kSNP3 Division of Disease Control
Date First Use
2022/03
Updated
NA
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
NA
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
kSNP3 can use multiple algorithms (maximum-likelihood, parsimony, neighbor-joining) to infer phylogenetic trees from genomes.
Purpose Desc
Produces phylogenetic trees which are used to rule in or out bacteria such as N. meningitidis in outbreaks.
Updated Desc
NA
Identifying Info
NA
Data Training
NA
Data Input
NA
Data Output
NA
Vendor Name
NA
Vendor Type
NA
Vendor Desc
NA
Data 2022
Assembled nucleic acid sequences recovered from pathogen genomes using high throughput sequencing as input data to this tool
Vendor
None
Analysis Type
NA
Department of Health and Mental Hygiene
Year: 2022 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2022
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
Newborn Home Visiting Program Screening Tools
Date First Use
2015/09
Updated
NA
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
NA
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
The NHVP database uses an algorithm based on screening tool results to determine breastfeeding and maternal depression referrals.
Purpose Desc
Based on the results from the PHQ-9 and GAD-7 screening tools, a client is deemed eligible for a referral to the internal program social worker for follow up on maternal mental health services. In addition, a breastfeeding assessment tool also embedded within the NHVP database is used to make referrals to the program's lactation support services. Based on the client's score, there is a decision made around the intensity of lactation support services that the client will receive.
Updated Desc
NA
Identifying Info
NA
Data Training
NA
Data Input
NA
Data Output
NA
Vendor Name
NA
Vendor Type
NA
Vendor Desc
NA
Data 2022
Data comes from standard tools around depression (PHQ-9) and anxiety (GAD-7). The threshold for the depression screening is a score of 10 and above, and threshold for the anxiety screening tool is also a score of 10 or above. This triggers a referral with the client's consent to the program's social worker for services, unless there is an immediate need identified such as imminent harm to self or others. Data from the breastfeeding tool is obtained in a questionnaire completed through the lead home visitor or the International Board Certified Lactation Consultant (IBCLC's) observation around baby's readiness to feed, latching, parental comfort to feed the baby, observation of milk transfer from mother to baby and parental report of baby's stool/urine output frequency.
Vendor
The Newborn Home Visiting Program Database is a DOHMH DITT Developed in-house system. The breastfeeding screening tool was developed within the program by a certified nurse.
Analysis Type
NA
Department of Health and Mental Hygiene
Year: 2022 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2022
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
Pangolin Division of Disease Control
Date First Use
2021/07
Updated
NA
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
NA
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
Pangolin uses a combination of several methods, including random forest tree classification methods, and maximum parsimony to assign lineage names to SARS-CoV-2 genomic sequences to bin sequences that are more likely to be similar.
Purpose Desc
Assigns SARS-CoV-2 viral sequences obtained from clinical specimens to known lineages. Used to determine circulating lineages in NYC.
Updated Desc
NA
Identifying Info
NA
Data Training
NA
Data Input
NA
Data Output
NA
Vendor Name
NA
Vendor Type
NA
Vendor Desc
NA
Data 2022
Nucleic acid sequences recovered from pathogen genomes using high throughput sequencing. The ML model is trained using the sequence data collected from around the world to generate variant nomenclature. Input data are designated with a variant name by maximum parsimony and random forest tree classification methods.
Vendor
None
Analysis Type
NA
Department of Health and Mental Hygiene
Year: 2022 • Agency: Department of Health and Mental Hygiene • Department: NA
Year
2022
Agency
Department of Health and Mental Hygiene
Department
NA
Tool Name
PHYLOViZ Division of Disease Control
Date First Use
2017/10
Updated
NA
Purpose Type
NA
Computation Type
NA
Autonomy
NA
Frequency
NA
Population Type
NA
Population Type Individual
NA
Population Type Other
NA
Website
NA
Tool Desc
For representing the possible evolutionary relationships between strains, PHYLOViZ uses the goeBURST algorithm, a refinement of eBURST algorithm by Feil et al., and its expansion to generate a complete minimum spanning tree (MST)
Purpose Desc
Used to generate the minimum spanning tree relationships which are used to rule in or out Legionella strains in outbreaks.
Updated Desc
NA
Identifying Info
NA
Data Training
NA
Data Input
NA
Data Output
NA
Vendor Name
NA
Vendor Type
NA
Vendor Desc
NA
Data 2022
Assembled nucleic acid sequences recovered from pathogen genomes using high throughput sequencing and qualitative data to represent sample and patient data.
Vendor
None
Analysis Type
NA