Impact Evaluations in Data Poor Settings: The Case of Stress-Tolerant Rice Varieties in Bangladesh
Impact evaluations of new technologies are critical to assessing and improving investment in national and international development goals. Yet many of these technologies are introduced and promoted at times and in places that lack the necessary data to conduct a strongly identified impact evaluation. We present a new method that combines remotely sensed Earth observation (EO) data, recent advances in machine learning, and socioeconomic survey data so as to allow researchers to conduct impact evaluations of a certain class of technologies when traditional economic data is missing. To demonstrate our approach, we study stress tolerant rice varieties (STRVs) that were introduced in Bangladesh more than a decade ago. Using 20 years of EO data on rice production and flooding, we fail to replicate existing RCT and field trial evidence of STRV effectiveness. We validate this failure to replicate with administrative and household panel data as well as conduct Monte Carlo simulations to test the sensitivity to mismeasurement of past evidence on the effectiveness of STRVs. Our findings speak to conducting large scale, long-term impact evaluations to verify external validity of small scale experimental data while also laying out a path for researchers to conduct similar evaluations in other data poor settings.
Code (0)
등록된 구현이 없습니다.
Tasks
Earth ObservationSimilar Papers 제목 키워드 기반
Position: Graph Learning Will Lose Relevance Due To Poor Benchmarks
While machine learning on graphs has demonstrated promise in drug design and molecular property prediction, significant benchmarking challenges hinder its further progress and relevance. Current benchmarking practices of…
BenchmarkingCombinatorial OptimizationDrug DesignGraph Learning+3Impacts of Racial Bias in Historical Training Data for News AI
AI technologies have rapidly moved into business and research applications that involve large text corpora, including computational journalism research and newsroom settings. These models, trained on extant data from var…
Reassessing Evaluation Practices in Visual Question Answering: A Case Study on Out-of-Distribution Generalization
Vision-and-language (V&L) models pretrained on large-scale multimodal data have demonstrated strong performance on various tasks such as image captioning and visual question answering (VQA). The quality of such models is…
Image CaptioningOut-of-Distribution GeneralizationQuestion AnsweringText Generation+2Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse
Chain-of-thought (CoT) prompting has become a widely used strategy for improving large language and multimodal model performance. However, it is still an open question under which settings CoT systematically reduces perf…
The Digital Transformation in Health: How AI Can Improve the Performance of Health Systems
Mobile health has the potential to revolutionize health care delivery and patient engagement. In this work, we discuss how integrating Artificial Intelligence into digital health applications-focused on supply chain, pat…
Management