The goal of washopenresearch is to provide an overview of open research data related to Water Sanitation and Hygiene (WASH). The current version contains seven datasets from the following sources:
-
washdev: Open access journal Journal of Water, Sanitation and Hygiene for Development -
ws: Journal Water Supply -
jwh: Journal Journal of Water and Health -
aqua: Journal AQUA - Water Infrastructure, Ecosystems and Society -
uncnewsletter: Research section of the newsletter North Carolina Water News -
ploswater: Open access journal PLOS Water -
datapapers: WASH-related data papers in seven dedicated data journals (Scientific Data, Data in Brief, Gates Open Research, F1000Research, GigaScience, GigaByte, and Data), harvested from Crossref and Europe PMC

Installation
You can install the development version of washopenresearch from GitHub with:
# install.packages("devtools")
devtools::install_github("openwashdata/washopenresearch")Alternatively, you can download the individual datasets as a CSV or XLSX file from the table below.
| dataset | CSV | XLSX |
|---|---|---|
| washdev | Download CSV | Download XLSX |
| uncnewsletter | Download CSV | Download XLSX |
| ploswater | Download CSV | Download XLSX |
| datapapers | Download CSV | Download XLSX |
| ws | Download CSV | Download XLSX |
| jwh | Download CSV | Download XLSX |
| aqua | Download CSV | Download XLSX |
Data
The package provides access to seven datasets washdev, ws, jwh, aqua, uncnewsletter, ploswater, and datapapers. Each dataset collects information on scientific articles about (1) article metadata (e.g. title, first author, correspondence author), (2) supplementary material information, (3) data availability statement or linked data repository, and (4) semantic information (e.g. keywords or abstract).
Author email addresses were removed in version 0.4.0. They are personal data, and the research questions this package serves use author country, data availability statement type, keywords and supplementary counts rather than contact details. Author names, affiliations and ORCIDs are kept, since ORCIDs are persistent public identifiers built for attribution.
Coverage and updates
Five datasets are updated monthly. Each data release adds the journal issues published since the last one. uncnewsletter and datapapers are frozen: they are still built with every release but no longer updated. The table shows what each dataset covers in this version.
| Dataset | Updates | Articles | Years | Latest issue |
|---|---|---|---|---|
washdev |
monthly | 1,173 | 2011 to 2026 | Vol. 16 Issue 6 |
ws |
monthly | 4,884 | 2001 to 2026 | Vol. 26 Issue 6 |
jwh |
monthly | 2,013 | 2003 to 2026 | Vol. 24 Issue 6 |
aqua |
monthly | 1,819 | 1998 to 2026 | Vol. 75 Issue 6 |
ploswater |
monthly | 434 | 2022 to 2026 | Vol. 5 Issue 7, published up to 2026-07-06 |
uncnewsletter |
frozen | 173 | 2020 to 2023 | newsletter ceased in May 2024 |
datapapers |
frozen | 8 | 2018 to 2025 | harvested on 23 July 2026 |
washdev
The dataset washdev contains data on open access articles of the Journal of Water, Sanitation & Hygiene for Development (Vol. 1 Issue 1 to Vol. 16 Issue 6). It has 1173 observations from March 2011 to 2026.
washdev |>
head(3) |>
gt::gt() |>
gt::as_raw_html()| paperid | volume | issue | paper_url | journal | title | published_year | is_supp | num_supp | supp_file_type | supp_url | num_authors | first_author_name | first_author_affiliation | first_author_affiliation_country | first_author_orcid | correspondence_author_name | correspondence_author_affiliation | correspondence_author_affiliation_country | correspondence_author_orcid | has_das | das | das_type | das_repo_url | keywords | doi | url_source |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
For an overview of the variable names, see the following table.
| variable_name | variable_type | description |
|---|---|---|
| paperid | integer | ID number of the paper on the journal website |
| volume | integer | Volume number of the journal |
| issue | integer | Issue number of the journal |
| paper_url | character | Official website url of the paper |
| journal | character | Full name of the journal |
| title | character | Title of the paper |
| published_year | integer | Year of publication |
| is_supp | logical | Whether the paper has supplementary materials |
| num_supp | integer | Number of supplementary material files |
| supp_file_type | character | File types of the supplementary materials separated by a semicolon when there are multiple |
| supp_url | character | Website urls of the supplementary materials separated by a semicolon when there are multiple |
| num_authors | integer | Number of the authors |
| first_author_name | character | Name of the first author |
| first_author_affiliation | character | Academic affiliation of the first author |
| first_author_affiliation_country | character | Country of the first author parsed from first_author_affiliation variable encoded with United Nations names |
| first_author_orcid | character | ORCID of the first author |
| correspondence_author_name | character | Name of the correspondence author |
| correspondence_author_affiliation | character | Academic affiliation of the correspondence author |
| correspondence_author_affiliation_country | character | Country of the correspondence author parsed from correspondence_author_affiliation variable encoded with United Nations names |
| correspondence_author_orcid | character | ORCID of the correspondence author |
| has_das | logical | Whether the paper has a data availability statement |
| das | character | Original data availability statement of the paper. NA if it does not have a data availability statement. |
| das_type | factor | Type of the data availability statement including “in paper”(data in full paper scope like supplementary material or appendix or main content) “on request”(data available on request to the authors) “available in online repository”(data is shared in a public online repository) “not shareable”(data is not shareable). NA if it does not have a data availability statement. |
| das_repo_url | character | Website urls of the data if the relevant data of the paper is shared on a public repository separated by a semicolon when there are multiple |
| keywords | character | Keywords of the paper separated by a semicolon |
| url_source | character | Publisher website of the paper |
| doi | character | DOI of the paper. Collected by the R scraper for recent articles and backfilled via Crossref for legacy rows (issue #20); NA where no Crossref match was found (see data-raw/washdev-doi-review.csv) |
ws
The dataset ws contains data on all articles of the journal Water Supply, collected with the same R scraper as washdev and sharing its schema. It has 4884 observations from 2001 to 2026, the longest run in the package. Water Supply does not enforce a data availability policy, so most articles carry no statement at all; where one exists, das_type records what it claims.
ws |>
head(3) |>
gt::gt() |>
gt::as_raw_html()| paperid | volume | issue | paper_url | journal | title | published_year | is_supp | num_supp | supp_file_type | supp_url | num_authors | first_author_name | first_author_affiliation | first_author_affiliation_country | first_author_orcid | correspondence_author_name | correspondence_author_affiliation | correspondence_author_affiliation_country | correspondence_author_orcid | has_das | das | das_type | das_repo_url | keywords | doi | url_source |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
For an overview of the variable descriptions, see the following table.
| variable_name | variable_type | description |
|---|---|---|
| paperid | integer | ID number of the paper on the journal website |
| volume | integer | Volume number of the journal |
| issue | character | Issue number of the journal, as text because combined issues (for example “1-2”) occur |
| paper_url | character | Official website url of the paper |
| journal | character | Full name of the journal |
| title | character | Title of the paper |
| published_year | numeric | Year of publication |
| is_supp | logical | Whether the paper has supplementary materials |
| num_supp | integer | Number of supplementary material files |
| supp_file_type | character | File types of the supplementary materials separated by a semicolon when there are multiple |
| supp_url | character | Website urls of the supplementary materials separated by a semicolon when there are multiple |
| num_authors | integer | Number of the authors |
| first_author_name | character | Name of the first author |
| first_author_affiliation | character | Academic affiliation of the first author |
| first_author_affiliation_country | character | Country of the first author parsed from first_author_affiliation variable encoded with United Nations names |
| first_author_orcid | character | ORCID of the first author |
| correspondence_author_name | character | Name of the correspondence author |
| correspondence_author_affiliation | character | Academic affiliation of the correspondence author |
| correspondence_author_affiliation_country | character | Country of the correspondence author parsed from correspondence_author_affiliation variable encoded with United Nations names |
| correspondence_author_orcid | character | ORCID of the correspondence author |
| has_das | logical | Whether the paper has a data availability statement |
| das | character | Original data availability statement of the paper. NA if it does not have a data availability statement. |
| das_type | factor | Type of the data availability statement including “in paper”(data in full paper scope like supplementary material or appendix or main content) “on request”(data available on request to the authors) “available in online repository”(data is shared in a public online repository) “not shareable”(data is not shareable). NA if it does not have a data availability statement. |
| das_repo_url | character | Website urls of the data if the relevant data of the paper is shared on a public repository separated by a semicolon when there are multiple |
| keywords | character | Keywords of the paper separated by a semicolon |
| doi | character | DOI of the paper. Collected by the R scraper; populated for every row |
| url_source | character | Publisher website of the paper |
jwh
The dataset jwh contains data on all articles of the Journal of Water and Health, collected with the same R scraper as washdev and sharing its schema. It has 2013 observations from 2003 to 2026. Like Water Supply, it does not enforce a data availability policy.
jwh |>
head(3) |>
gt::gt() |>
gt::as_raw_html()| paperid | volume | issue | paper_url | journal | title | published_year | is_supp | num_supp | supp_file_type | supp_url | num_authors | first_author_name | first_author_affiliation | first_author_affiliation_country | first_author_orcid | correspondence_author_name | correspondence_author_affiliation | correspondence_author_affiliation_country | correspondence_author_orcid | has_das | das | das_type | das_repo_url | keywords | doi | url_source |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
For an overview of the variable descriptions, see the following table.
| variable_name | variable_type | description |
|---|---|---|
| paperid | integer | ID number of the paper on the journal website |
| volume | integer | Volume number of the journal |
| issue | character | Issue number of the journal, as text because combined issues (for example “1-2”) occur |
| paper_url | character | Official website url of the paper |
| journal | character | Full name of the journal |
| title | character | Title of the paper |
| published_year | numeric | Year of publication |
| is_supp | logical | Whether the paper has supplementary materials |
| num_supp | integer | Number of supplementary material files |
| supp_file_type | character | File types of the supplementary materials separated by a semicolon when there are multiple |
| supp_url | character | Website urls of the supplementary materials separated by a semicolon when there are multiple |
| num_authors | integer | Number of the authors |
| first_author_name | character | Name of the first author |
| first_author_affiliation | character | Academic affiliation of the first author |
| first_author_affiliation_country | character | Country of the first author parsed from first_author_affiliation variable encoded with United Nations names |
| first_author_orcid | character | ORCID of the first author |
| correspondence_author_name | character | Name of the correspondence author |
| correspondence_author_affiliation | character | Academic affiliation of the correspondence author |
| correspondence_author_affiliation_country | character | Country of the correspondence author parsed from correspondence_author_affiliation variable encoded with United Nations names |
| correspondence_author_orcid | character | ORCID of the correspondence author |
| has_das | logical | Whether the paper has a data availability statement |
| das | character | Original data availability statement of the paper. NA if it does not have a data availability statement. |
| das_type | factor | Type of the data availability statement including “in paper”(data in full paper scope like supplementary material or appendix or main content) “on request”(data available on request to the authors) “available in online repository”(data is shared in a public online repository) “not shareable”(data is not shareable). NA if it does not have a data availability statement. |
| das_repo_url | character | Website urls of the data if the relevant data of the paper is shared on a public repository separated by a semicolon when there are multiple |
| keywords | character | Keywords of the paper separated by a semicolon |
| doi | character | DOI of the paper. Collected by the R scraper; populated for every row |
| url_source | character | Publisher website of the paper |
aqua
The dataset aqua contains data on all articles of the journal AQUA - Water Infrastructure, Ecosystems and Society in the journal’s online archive, collected with the same R scraper as washdev and sharing its schema. It has 1819 observations from 1998 to 2026, of which 539 carry a data availability statement. The journal changed its title during this period; the dataset uses the current title for every year.
aqua |>
head(3) |>
gt::gt() |>
gt::as_raw_html()| paperid | volume | issue | paper_url | journal | title | published_year | is_supp | num_supp | supp_file_type | supp_url | num_authors | first_author_name | first_author_affiliation | first_author_affiliation_country | first_author_orcid | correspondence_author_name | correspondence_author_affiliation | correspondence_author_affiliation_country | correspondence_author_orcid | has_das | das | das_type | das_repo_url | keywords | doi | url_source |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
For an overview of the variable descriptions, see the following table.
| variable_name | variable_type | description |
|---|---|---|
| paperid | integer | ID number of the paper on the journal website |
| volume | integer | Volume number of the journal |
| issue | character | Issue number of the journal, as text because combined issues (for example “1-2”) occur |
| paper_url | character | Official website url of the paper |
| journal | character | Full name of the journal |
| title | character | Title of the paper |
| published_year | numeric | Year of publication |
| is_supp | logical | Whether the paper has supplementary materials |
| num_supp | integer | Number of supplementary material files |
| supp_file_type | character | File types of the supplementary materials separated by a semicolon when there are multiple |
| supp_url | character | Website urls of the supplementary materials separated by a semicolon when there are multiple |
| num_authors | integer | Number of the authors |
| first_author_name | character | Name of the first author |
| first_author_affiliation | character | Academic affiliation of the first author |
| first_author_affiliation_country | character | Country of the first author parsed from first_author_affiliation variable encoded with United Nations names |
| first_author_orcid | character | ORCID of the first author |
| correspondence_author_name | character | Name of the correspondence author |
| correspondence_author_affiliation | character | Academic affiliation of the correspondence author |
| correspondence_author_affiliation_country | character | Country of the correspondence author parsed from correspondence_author_affiliation variable encoded with United Nations names |
| correspondence_author_orcid | character | ORCID of the correspondence author |
| has_das | logical | Whether the paper has a data availability statement |
| das | character | Original data availability statement of the paper. NA if it does not have a data availability statement. |
| das_type | factor | Type of the data availability statement including “in paper”(data in full paper scope like supplementary material or appendix or main content) “on request”(data available on request to the authors) “available in online repository”(data is shared in a public online repository) “not shareable”(data is not shareable). NA if it does not have a data availability statement. |
| das_repo_url | character | Website urls of the data if the relevant data of the paper is shared on a public repository separated by a semicolon when there are multiple |
| keywords | character | Keywords of the paper separated by a semicolon |
| doi | character | DOI of the paper. Collected by the R scraper; populated for every row |
| url_source | character | Publisher website of the paper |
uncnewsletter
The dataset uncnewsletter contains data on a curated list of articles published at the Research section of the newsletter North Carolina Water News. It has 173 observations from 2020 to 2023. The newsletter ceased publication in May 2024, so this dataset is a frozen source.
uncnewsletter |>
head(3) |>
gt::gt() |>
gt::as_raw_html()| paperid | issue_url | paper_url | url_source | journal | title | published_year | is_supp | num_supp | supp_file_type | supp_url | num_authors | first_author_name | first_author_affiliation | first_author_affiliation_country | first_author_orcid | correspondence_author_name | correspondence_author_affiliation | correspondence_author_affiliation_country | correspondence_author_orcid | has_das | das | das_type | das_repo_url | citations | keywords | doi |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
For an overview of the variable descriptions, see the following table.
| variable_name | variable_type | description |
|---|---|---|
| paperid | integer | ID number of the paper on the journal website |
| issue_url | character | URL of the newsletter issue that featured the paper |
| paper_url | character | Official website url of the paper |
| url_source | character | Publisher website of the paper |
| journal | character | Full name of the journal |
| title | character | Title of the paper |
| published_year | integer | Year of publication |
| is_supp | logical | Whether the paper has supplementary materials |
| num_supp | integer | Number of supplementary material files |
| supp_file_type | character | File types of the supplementary materials separated by a semicolon when there are multiple |
| supp_url | character | Website urls of the supplementary materials separated by a semicolon when there are multiple |
| num_authors | integer | Number of the authors |
| first_author_name | character | Name of the first author |
| first_author_affiliation | character | Academic affiliation of the first author |
| first_author_affiliation_country | character | Country of the first author directly parsed from first_author_affiliation variable encoded with United Nation names |
| first_author_orcid | character | ORCID of the first author |
| correspondence_author_name | character | Name of the correspondence author |
| correspondence_author_affiliation | character | Academic affiliation of the correspondence author |
| correspondence_author_affiliation_country | character | Country or region of the correspondence author directly parsed from correspondence_author_affiliation variable encoded with United Nation names |
| correspondence_author_orcid | character | ORCID of the correspondence author |
| has_das | logical | Whether the paper has a data availability statement |
| das | character | Original data availability statement of the paper. NA if it does not have a data availability statement. |
| das_type | factor | Type of the data availability statement including “in paper”(data in full paper scope like supplementary material or appendix or main content) “on request”(data available on request to the authors) “available in online repository”(data is shared in a public online repository) “not shareable”(data is not shareable). NA if it does not have a data availability statement. |
| das_repo_url | character | Website urls of the data if the relevant data of the paper is shared on a public repository separated by a semicolon when there are multiple |
| citations | numeric | Number of citations of the paper as entered by the annotators during the manual collection in January 2024 or earlier; the source of the count is not documented. NA where no value was entered. |
| keywords | character | Keywords of the paper separated by a semicolon |
| doi | character | DOI of the paper backfilled via a Crossref title search (issue #20); NA where no match cleared the title-similarity threshold (see data-raw/uncnewsletter-doi-review.csv) |
ploswater
The dataset ploswater contains data on all articles of the journal PLOS Water from its first volume (2022) onward, collected through the public PLOS API rather than web scraping. It has 434 observations. Data availability statements are mandatory at PLOS, so the interesting variation lies in das_type, das_repo_url, and das_repo_name, which describe how and where the data behind each article is stored. All article types are included; use article_type to restrict to research articles.
ploswater |>
head(3) |>
gt::gt() |>
gt::as_raw_html()| paperid | volume | issue | paper_url | journal | title | published_year | is_supp | num_supp | supp_file_type | supp_url | num_authors | first_author_name | first_author_affiliation | first_author_affiliation_country | first_author_orcid | correspondence_author_name | correspondence_author_affiliation | correspondence_author_affiliation_country | correspondence_author_orcid | has_das | das | das_type | das_repo_url | das_repo_name | keywords | url_source | doi | article_type | publication_date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
For an overview of the variable descriptions, see the following table.
| variable_name | variable_type | description |
|---|---|---|
| paperid | character | DOI of the paper; identical to the doi variable |
| volume | integer | Volume number of the journal; volume 1 is 2022 |
| issue | integer | Issue number of the journal |
| paper_url | character | Official website url of the paper |
| journal | character | Full name of the journal |
| title | character | Title of the paper |
| published_year | integer | Year of publication |
| is_supp | logical | Whether the paper has supplementary materials |
| num_supp | integer | Number of supplementary material files |
| supp_file_type | character | File types of the supplementary materials separated by a semicolon when there are multiple |
| supp_url | character | Website urls of the supplementary materials separated by a semicolon when there are multiple. Stable download endpoints that do not expire |
| num_authors | integer | Number of the authors |
| first_author_name | character | Name of the first author |
| first_author_affiliation | character | Academic affiliation of the first author |
| first_author_affiliation_country | character | Country of the first author parsed from first_author_affiliation encoded with United Nations names |
| first_author_orcid | character | ORCID of the first author |
| correspondence_author_name | character | Name of the correspondence author |
| correspondence_author_affiliation | character | Academic affiliation of the correspondence author |
| correspondence_author_affiliation_country | character | Country of the correspondence author parsed from correspondence_author_affiliation encoded with United Nations names |
| correspondence_author_orcid | character | ORCID of the correspondence author |
| has_das | logical | Whether the paper has a data availability statement |
| das | character | Original data availability statement of the paper. NA if it does not have a data availability statement |
| das_type | factor | Type of the data availability statement including “available in online repository”(data is shared in a public online repository) “in paper”(data in full paper scope like supplementary material or appendix or main content) “on request”(data available on request to the authors) “not shareable”(data is not shareable) “no data generated”(the study produced no datasets). NA if it does not have a data availability statement or no classification rule matched |
| das_repo_url | character | Website urls and dataset DOIs mentioned in the data availability statement separated by a semicolon when there are multiple |
| das_repo_name | character | Recognized data repositories behind das_repo_url (e.g. zenodo dryad figshare osf github dataverse) separated by a semicolon when there are multiple |
| keywords | character | Subject terms of the paper from the PLOS search API separated by a semicolon. PLOS Water articles carry no author keywords in their XML |
| url_source | character | Publisher website of the paper |
| doi | character | DOI of the paper |
| article_type | character | Article type e.g. Research Article or Opinion or Review |
| publication_date | date | Date of publication (ISO 8601) |
datapapers
The dataset datapapers contains WASH-related data papers published in seven dedicated data journals, identified from Crossref and Europe PMC metadata and screened for relevance (see data-raw/README.md for the pipeline). It has 8 observations. The dataset is frozen: it was harvested once and its screening is closed. Because a data paper exists to describe a shared dataset, data_repo_url and data_repo take the role that the data availability statement variables play in the other two datasets.
datapapers |>
head(3) |>
gt::gt() |>
gt::as_raw_html()| paperid | doi | paper_url | url_source | journal | title | published_year | num_authors | first_author_name | first_author_affiliation | first_author_affiliation_country | data_repo_url | data_repo | license | related_paper_doi | abstract | query_term | retrieval_date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
For an overview of the variable descriptions, see the following table.
| variable_name | variable_type | description |
|---|---|---|
| paperid | integer | ID number of the paper within this dataset |
| doi | character | DOI of the data paper |
| paper_url | character | Official url of the paper (DOI resolver link) |
| url_source | character | Publisher website of the paper |
| journal | character | Full name of the journal |
| title | character | Title of the paper |
| published_year | integer | Year of publication |
| num_authors | integer | Number of the authors |
| first_author_name | character | Name of the first author |
| first_author_affiliation | character | Academic affiliation of the first author |
| first_author_affiliation_country | character | Country of the first author parsed from first_author_affiliation variable encoded with United Nations names |
| data_repo_url | character | Website urls of the repository holding the dataset the paper describes separated by a semicolon when there are multiple |
| data_repo | character | Name of the data repository (e.g. Zenodo Dryad Figshare OSF Dataverse) parsed from data_repo_url |
| license | character | License url of the paper from Crossref metadata |
| related_paper_doi | character | DOI of a linked research article if any separated by a semicolon when there are multiple |
| abstract | character | Abstract of the paper as provided by the metadata source |
| query_term | character | WASH search term(s) that retrieved the paper separated by a semicolon |
| retrieval_date | date | Date the paper metadata was harvested from the API |
Example
washdev
- What are the top 10 countries(or regions) the first authors from in the Journal of Water, Sanitation and Hygiene for Development?
library(washopenresearch)
washdev |>
filter(!is.na(first_author_affiliation_country)) |>
group_by(first_author_affiliation_country) |>
summarise(count=n()) |>
arrange(desc(count)) |>
head(10) |>
ggplot() +
geom_col(aes(x = reorder(first_author_affiliation_country, count),
y = count)) +
labs(title = "Top 10 countries of first author",
subtitle = "in the Journal of Water, Sanitation and Hygiene for Development",
x = "First Author Country", y = "Count") +
scale_x_discrete(labels = scales::label_wrap(15))+
coord_flip() +
theme_classic()
- What are the top choices of keywords in WASH Dev?
Each publication may provide a list of keywords, typically 5-7, to summarize the topics of the article. Here we compile all keywords and calculate their frequency to be used.
keywords_freq <- washdev$keywords |>
str_split("; ") |>
unlist() |>
str_to_lower() |>
table() |>
as.data.frame() |>
as_tibble() |>
arrange(desc(Freq))
# Top 20 keywords
ggplot(data = head(keywords_freq, 20)) +
geom_bar(aes(x = reorder(Var1, Freq), y=Freq), stat = "identity") +
coord_flip() +
labs(title = "Top 20 Keywords in WASH Dev Journal", x = "Keywords", y = "Count") +
theme_bw()
uncnewsletter
- What are the top 10 source websites of the publications selected by the newsletter?
uncnewsletter |>
group_by(url_source) |>
summarise(count=n()) |>
arrange(desc(count)) |>
head(10) |>
ggplot() +
geom_col(aes(x = reorder(url_source, count),
y = count)) +
labs(title = "Top 10 publication websites",
subtitle = "in the selection of North Carolina Water News",
x = "Website URL", y = "Count") +
scale_x_discrete(labels = scales::label_wrap(15))+
coord_flip() +
theme_classic()
datapapers
- How many papers per journal, and how many resolve to a data repository?
datapapers |>
group_by(journal) |>
summarise(papers = n(),
with_repository_link = sum(!is.na(data_repo_url))) |>
arrange(desc(papers)) |>
knitr::kable()| journal | papers | with_repository_link |
|---|---|---|
| Scientific Data | 5 | 5 |
| Data | 2 | 1 |
| GigaByte | 1 | 1 |
Method
We describe the raw data collection procedure of each dataset in this section. The collection is scripted in R; the scripts live in data-raw/ and the run order is documented in data-raw/README.md.
washdev, ws, jwh and aqua
All four IWA journals are collected by the same R scraper. First, each publication link is scraped by iterating the table of contents of all volumes. This step delivers a table containing the paper ID, volume number, issue number, publication url, journal title, publication title, and published year. Then, for each publication, the remaining variables are retrieved from the article’s html using that url, rule-based, to find the relevant fields (for example supplementary materials) and extract the value.
All four journals are scraped with data-raw/iwa_scraping.R. Because they share one schema, the cleaning that follows is shared too: process_iwa_journal() in data-raw/helpers.R handles encoding repair, column harmonisation, country standardisation and the multi-value fields, while the data availability statement mapping and the review files stay per-journal.
Statements that no rule maps keep their full text and are listed in a review file per journal (data-raw/*-das-review.csv), so nothing is silently reclassified. The same applies to affiliations whose country could not be standardised (data-raw/*-country-review.csv).
datapapers
The collection of datapapers is fully scripted in R. Crossref is queried by journal ISSN and Europe PMC by journal name (for the F1000-platform journals) with a fixed list of WASH search terms; the harvest is committed as a raw snapshot with the retrieval date and matching query terms recorded per row. Relevance screening and country corrections are captured in committed CSV decision sheets keyed on DOI, so the pipeline runs end-to-end non-interactively. See data-raw/README.md for the run order.
uncnewsletter
The collection of uncnewsletter is a combination of web scraping and manual annotation. We first use the newsletter archive to scrape all publication website links. The code can be found at inst/python/uncnewsletter_scraping.py, which was removed from the package when the Python tooling went (#17); it is recoverable from the git history. Two annotators worked on the manual extraction of the needed variables on these publications. For each publication, an annotator follows the guide to fill in the value on an collaborative spreadsheet. The guide is converted into the data dictionary for this dataset.
License
Data are available as CC-BY.
Citation
Please cite this package using:
citation("washopenresearch")
#> To cite package 'washopenresearch' in publications use:
#>
#> Zhong M, Luz L, Schöbitz L, Dubey Y (2026). "washopenresearch:
#> Dataset about open research data information in Water, Sanitation,
#> and Hygiene." doi:10.5281/zenodo.11185699
#> <https://doi.org/10.5281/zenodo.11185699>.
#> <https://github.com/openwashdata/washopenresearch>.
#>
#> A BibTeX entry for LaTeX users is
#>
#> @Misc{zhong_etall:2026,
#> title = {washopenresearch: Dataset about open research data information in Water, Sanitation, and Hygiene},
#> author = {Mian Zhong and Ludwig Luz and Lars Schöbitz and Yash Dubey},
#> year = {2026},
#> doi = {10.5281/zenodo.11185699},
#> url = {https://github.com/openwashdata/washopenresearch},
#> abstract = {The goal of washopenresearch is to provide an overview of open research data related to Water Sanitation and Hygiene (WASH). The package provides access to seven datasets: `washdev`, `ws`, `jwh`, `aqua`, `uncnewsletter`, `ploswater`, and `datapapers`. Each dataset collects information on scientific articles about (1) article metadata (e.g. title, first author, correspondence author), (2) supplementary material information, (3) data availability statement, and (4) semantic information (e.g. keywords).},
#> keywords = {open-data,open-research-data,open-science,openwashdata,sanitation,wash},
#> version = {0.5.0},
#> }