Inquiry
Dear DOI Open Data Coordinator,
Together with a group known as the ***, I’ve been working through the new Evictions dataset published by DOI for pending, scheduled, and executed evictions starting in 2017. I have a number of questions regarding the dataset that I would greatly appreciate any insight into:
1. We are interested in de-duplicating the dataset to measure aggregate evictions throughout the five boroughs , and there are a number of observations that have a) identical index & docket numbers or b) identical marshal name, address, date, and unit number. Do either the former or the latter indicate duplicate observations? In the former situation, it would make sense that a unique index & docket combination would refer to a unique eviction, as the index # seems to be a unique ID provided by NY Housing Court, and docket # seems to be a unique ID assigned by the marshal. However, there are cases of duplicate index & docket combination assigned to different eviction addresses, unit #s, etc. In the latter situation, we are not sure whether we can rule out the possibility of these duplicates referring to, for example, same-day evictions in different units in the same building (where a marshal might have just recorded one unit # for ease). Please advise on how we should be thinking of de-duplicating this dataset.
2. Does this dataset constitute the universe of eviction judgments in the five boroughs from the NY Housing Courts? In other words, are there examples where eviction judgments are determined for a certain tenant that goes through NY Housing Court where a Marshal’s eviction would not automatically be scheduled and thus recorded in this dataset?
3. Relatedly, in the ‘Schedule_Status’ column of the dataset, all ~38k observations are listed as ‘Scheduled’ (in the data dictionary the column is described as either having values ‘Scheduled’, ‘Pending’, or ‘Confirmed). Assuming this is a mistake, when can we expect an update to the data in this column?
4. The data come with addresses of variable accuracy and without geospatial information, which is in violation of Local Law 108 of 2015, which requires, “every public data set containing address information to utilize a standard field layout and presentation of address information and include corresponding community district and geospatial reference data.” When can standardized addresses and geospatial reference data be expected to be appended to the dataset?
Thanks very much for your engagement with the above questions – we are looking forward to hearing back.
Best,
Assigned Agencys Response
Thank you for contacting NYC Open Data Help Desk. Please see agency's response below:
1. We are interested in de-duplicating the dataset to measure aggregate evictions throughout the five boroughs, and there are a number of observations that have a) identical index & docket numbers or b) identical marshal name, address, date, and unit number. Do either the former or the latter indicate duplicate observations? In the former situation, it would make sense that a unique index & docket combination would refer to a unique eviction, as the index # seems to be a unique ID provided by NY Housing Court, and docket # seems to be a unique ID assigned by the marshal. However, there are cases of duplicate index & docket combination assigned to different eviction addresses, unit #s, etc. In the latter situation, we are not sure whether we can rule out the possibility of these duplicates referring to, for example, same-day evictions in different units in the same building (where a marshal might have just recorded one unit # for ease). Please advise on how we should be thinking of de-duplicating this dataset.
The court assigns each eviction/legal possession case a unique index number. Docket numbers are assigned by a marshal’s office but one case may carry different docket numbers. For example, should the landlord instruct the marshal to request separate warrants for each of the respondents named in the case, each respondent will be entered as a separate docket. Some marshals add a letter “b or “c” to the docket number, which is an indication that there is more than one warrant of eviction associated to one index number. Also, an index number may be linked to different floors within the same address (Example, 2 or 3 family home occupied by an entire family or by unidentified squatters. In sum, one index number, with one or more docket numbers, is one or more eviction(s).
Moreover, different county courts (Kings, Queens, Bronx, Manhattan and Richmond) may use the same index number used in other counties for different eviction cases. Some marshals may add the first letter of the corresponding county to indicate which borough it came from. In short, there can be one index number, for different addresses in different counties.
2. Does this dataset constitute the universe of eviction judgments in the five boroughs from the NY Housing Courts? In other words, are there examples where eviction judgments are determined for a certain tenant that goes through NY Housing Court where a Marshal’s eviction would not automatically be scheduled and thus recorded in this dataset?
This dataset will never constitute the entire universe of “eviction judgments.”
Petitioners may obtain a court ordered judgment of eviction and, through later negotiations among the parties, the matter is settled privately without a marshal ever having been contacted by the petitioner. These private negotiations/agreements between petitioner and respondent may take place at any time (before the marshal requests the warrant, before or after the marshal serves notice, after the marshal has scheduled an eviction…)
There was one marshal who was not using the program because according to the handbook, he had a low volume of eviction cases and he did not have to be computerized. However, the marshal is no longer doing evictions.
3. Relatedly, in the ‘Schedule Status’ column of the dataset, all ~38k observations are listed as ‘Scheduled’ (in the data dictionary the column is described as either having values ‘Scheduled’, ‘Pending’, or ‘Confirmed). Assuming this is a mistake, when can we expect an update to the data in this column?
This ‘Scheduled Status’ field should read either “Pending,” meaning not yet scheduled, or “Scheduled” meaning it’s on the marshal’s calendar. This is the way it should appear on open data. The actual scheduled date and pending evictions are not for publication due to safety reasons. The CPR has a “Confirmed” field which is an automatic confirmation, generated by the CPR system, to the marshals confirming that the next day’s eviction schedule was received by the system, no later than 4:00 pm. If they don’t receive this confirmation they do not have DOI’s approval to go forward with their evictions. This has nothing to do with public information, this is for DOI’s regulatory purposes.
4. The data come with addresses of variable accuracy and without geospatial information, which is in violation of Local Law 108 of 2015, which requires, “every public data set containing address information to utilize a standard field layout and presentation of address information and include corresponding community district and geospatial reference data.” When can standardized addresses and geospatial reference data be expected to be appended to the dataset?
There is some merit to this point. Not all datasets published in OpenData come with geospatial information. In fact, not all development teams if any are trained to include exactly what is in LL108. If this was a problem, the Open Data team would have stopped publishing any dataset and required us to make remediation steps before publishing as they are the gatekeepers and we need to adhere to certain standards. Knowing this now, I will know better going forward since we have the capability now as well. To address the present problem specific to this set, it seems clear that the requester is asking for house numbers and street numbers to be split so the dataset can be geocoded either on our side or their side. The desktop geocoder is available for the public from DCP to use, so they would need to run the dataset through that if they wanted everything. Because the law doesn’t say which reference data other than Community District, the vagueness does not help guide us to what else is needed other than having the requester self-service the rest through the desktop geocoder. It not that it isn’t available, it just won’t be easy to get ‘the rest’ of whatever the requester is seeking. So we would need the house number and street name split for us to geocode it and it is work I would assist you putting in an intake to secure this as a project enhancement. This would be the prime opportunity, if you want to proceed, to add more fields that DOI may need (e.g. case type, number of Order to show cause). I would suggest the public request revision of LL108 to be more specific on what kind of geospatial reference data. As far as I know, the address is a valid piece of information that is technically geospatial reference data already without more specifics on the definition. Here’s the actual LL108. https://www1.nyc.gov/assets/buildings/local_laws/ll108of2015.pdf
Thank you,
-NYC Department of Investigation (DOI)