Did you know that we have access to digital copies of the historical US census rolls? You can also find the digitized data at IPUMS. However, the data for people with disabilities is not great. It depends on the year, but those data have error rates on the order of 20% or higher. We have the digital census rolls, the data just doesn’t match them.
So, I created a non-install windows computer application that lets people identify disabled people on those digital census rolls. Complemented with machine learning, my goal is to improve the accuracy of historical records about people with disabilities. Historical and quantitative research about disabled populations is relatively thin. We can do better. If you have students who would benefit from this research experience, then do please let me know! I can approve your institution’s email domain and we can get started.
The application is really straightforward with basically two user-facing features.
1) Image Identification

Disability information is included in column 18 of the 1870 schedule 1 census form. Enumerators were instructed to enter full phrases to denote disability. So, check marks and tick marks are not valid entries. In the current version of the app, users simply press zero or one in order to label whether a disabled person is present on the form or not. Once labeled, users are immediately presented with the next image. The entire application is pretty lightweight so that there is minimal lag. Users can easily label one thousand images in an hour.
On that same screen users can zoom and pan as necessary. They can also flag images for administrative review. This might be necessary for images that aren’t schedule 1 forms, are entirely blank, display multiple pages, or are otherwise corrupted.
2) Statistics
The server tracks some basic stats for each user. Straightforwardly, the number of labels are tracked for each user and each institution. They’re also ranked for some implicit competition. Less straightforward is the estimated number of images labeled correctly. There is some fancy statistical work going on the background. If we knew exactly when individuals were right and wrong, then we wouldn’t have needed the images labeled in the first place! Rather, user labels are compared against what other users think conditional on how often they agree with the consensus. From that, we can back out the estimate of correct labels.

If you are an instructor, then you could require that students label some number of images correctly. This prevents button mashers and low conscientious students from getting credit for poor label quality. If you tell me who is in your class, then I can even send you a csv with your roster and performances. If you are a student, then maybe suggest this project to your teacher as a way of earning extra credit. OR, if you need service hours for one reason or another, then please put your admin in contact with me! Like I said, one can easily complete about a thousand correct labels in an hour. Fulfilling your service hour requirement while wearing pajamas from the comfort of your own home isn’t too shabby!
Please share this far and wide and let me know if you’d like to contribute. You can inquire at contact@disabilityrecordsproject.org
Amazing to see this!
LikeLike