Changelog
Source:NEWS.md
washopenresearch 0.5.0
This release makes the build of the datasets reproducible and prepares monthly data updates. The data of the six datasets of 0.4.0 changes only through the fixes listed below, and aqua joins as the seventh dataset. No new journal issues are added yet; the first monthly update follows as 0.6.0.
Coverage
Five datasets are updated monthly from now on. uncnewsletter and datapapers are frozen: still built with every release, no longer updated.
| Dataset | Updates | Articles | Years | Latest issue |
|---|---|---|---|---|
washdev |
monthly | 1,173 | 2011 to 2026 | Vol. 16 Issue 6 |
ws |
monthly | 4,884 | 2001 to 2026 | Vol. 26 Issue 6 |
jwh |
monthly | 2,013 | 2003 to 2026 | Vol. 24 Issue 6 |
aqua |
monthly | 1,819 | 1998 to 2026 | Vol. 75 Issue 6 |
ploswater |
monthly | 434 | 2022 to 2026 | Vol. 5 Issue 7, published up to 2026-07-06 |
uncnewsletter |
frozen | 173 | 2020 to 2023 | newsletter ceased in May 2024 |
datapapers |
frozen | 8 | 2018 to 2025 | harvested on 23 July 2026 |
Reproducible build
The datasets are built by a
targetspipeline from the raw snapshots and decision sheets committed indata-raw/, with the package versions pinned inrenv.lock. One command rebuilds everything:Rscript data-raw/build.R. Manual corrections moved from the code into decision sheets keyed on the article identifier.A validation gate checks every dataset before anything is written: documented columns, a unique key, no email addresses or credentials, statement text wherever a statement is flagged, UN country names, and no article lost since the last release.
A workflow on GitHub rebuilds the datasets on every push and pull request and fails when they differ from the committed data (
data-raw/check_reproducible.R).Zenodo files the release as a dataset in the openwashdata community, from the new
.zenodo.json. The releases 0.1.0, 0.3.0 and 0.4.0 were never archived there.
New features
-
New dataset
aqua: all 1,819 articles of the journal AQUA - Water Infrastructure, Ecosystems and Society in the journal’s online archive, 1998 to 2026, with the same columns aswsandjwh. 539 articles carry a data availability statement; 529 of these map to a shareddas_typeand the remaining 10 keep their full text and are listed indata-raw/aqua-das-review.csv.The 0.4.0 notes said AQUA needed a re-scrape because its statement text had never been captured. That was wrong. The statements were in the raw snapshot all along. The build read the snapshot with guessed column types, the journal’s first statement sits beyond the rows a guess looks at, so the statement column was read as logical and came back empty. The build now states every column type (#52).
Bug fixes
washdev: 33 correspondence author countries were three letter ISO codes (“IND”, “GBR”) in a column that otherwise holds United Nations country names. 32 are now the UN name. The 33rd, an affiliation in Taiwan, is missing, as it already was in the first author column: the UN names have no entry for it. The article is listed indata-raw/washdev-country-review.csv.ploswater: two articles appeared twice (10.1371/journal.pwat.0000024 and 10.1371/journal.pwat.0000520), each as two rows identical in every column. The dataset now has 434 rows, one per article. The cause was a search query paged without a sort order.Credentials are removed from the text of the datasets. Where a statement linked a restricted Zenodo record through an access token, the token parameter is now dropped and the plain record URL kept; up to v0.4.0 the URL kept the parameter with a placeholder value and did not resolve (
ploswater, one article). Where authors wrote a password into their statement, such as the login of an FTP site, the password is replaced by[removed]and the rest of the sentence kept (ploswaterandws, one statement each).ws: one data availability statement kept an author email address indas_type. The masking of addresses inside text, introduced in v0.4.0, skipped that column because it is a factor. The address is now masked there as it already was indas.
Changes
washdev,ws,jwh: the levels ofdas_typenow list the shared statement types first (“available in online repository”, “in paper”, “on request”), followed by the statements no rule mapped. Before, the three types sat scattered among the statements in alphabetical order. Values are unchanged; code that relies on the integer codes of the factor needs checking.Credentials are removed from the raw snapshots in
data-raw/, as they already were from the datasets. The download links of supplementary files inwashdev.csv,ws.csv,jwh.csvandaqua.csvlose their signed query and keep the file path (1,752 links of 1,532 articles, all expired since August 2026). Inploswater.csvthe access tokens in the URLs of one statement are dropped. A password that authors wrote into their statement is replaced by[removed](ploswater.csvandws.csv, one statement each). The scrapers apply the same rules to new articles. The datasets are unchanged. The git history was not rewritten, so commits up to and including v0.4.0 still hold these values.-
The raw snapshots (
washdev.csv,ws.csv,jwh.csv,aqua.csvandploswater.csv) no longer have the columnsfirst_author_emailandcorrespondence_author_email, and the scrapers no longer collect author email addresses. Contact addresses that some article pages print after the affiliation are masked in the same way as in the datasets (washdev.csv10 articles,aqua.csv4), and the scraper masks them in new articles. No other value in the snapshots changed, and the datasets are unchanged.The git history was not rewritten. Commits up to and including v0.4.0 still hold the removed values, and the archive of v0.0.1 on Zenodo holds the email columns of
washdev.csv. Addresses that authors wrote inside a data availability statement stay in the raw statement text, which is the published statement, and are masked in the datasets as before.
Data acquisition
The scraper of the IWA journals is incremental. A run lists the journal issues again from the newest known publication year, scrapes the issues that are missing from the raw snapshot, and reads the two most recent issues again to fetch articles that were added late. An issue that the site lists before its articles are online is asked for again on the next run. The functions that decide what a run fetches are tested without network.
The scraper stores an issue complete or not at all, and writes no row for an article whose page was not read. A page counts as read only when the site served the kind of page that was asked for. Its “Not Found” page loads completely too and was taken for an article without authors, statement or supplement, or for an issue without articles.
washdevis scraped by the same scraper asws,jwhandaqua(data-raw/iwa_scraping.R). The separatedata-raw/washdev_scraping.Rand the runnerdata-raw/run_iwa_scrapes.Rare removed. The old washdev scraper read a throttled page as “no more issues” and stored an article that failed to load as an empty row. The index column that the first Python scraper left indata-raw/washdev.csvis removed. The datasetwashdevis unchanged.The signed download links of supplementary files, which work for about three weeks, are written to
data-raw/private/, which is not committed.data-raw/suppfiles_download.Rreads them there.One command runs the monthly acquisition for every live source:
Rscript data-raw/update_sources.R. It runs the four IWA journals in sequence with a cooldown, then PLOS Water, each as its own R process. A source that fails does not stop the others, anddata-raw/update-log.csvrecords per journal and run how many journal issues and rows were added.The PLOS Water downloader can run incrementally. It failed on the second run, because it read the existing snapshot with guessed column types and could not combine the publication date with the new rows. It now pages the search in a fixed sort order and keeps one row per DOI. The two repeated rows are removed from
data-raw/ploswater.csv; the dataset has been without them since the fix listed above. The unsorted paging had also skipped two articles (10.1371/journal.pwat.0000223 and 10.1371/journal.pwat.0000302). They enter the dataset with the next data update.A workflow opens an issue on the 5th of each month with the checklist of the monthly data update (
data-raw/monthly-data-update.md).
Documentation
The help pages and the README say which datasets are updated monthly and which are frozen, and the README has a coverage table that is computed from the datasets.
uncnewsletter: thecitationscolumn is now described in the help page and the data dictionary. It has been in the dataset since its first release without documentation. The source of the counts is not recorded.
washopenresearch 0.4.0
Breaking changes
The scraped author email addresses are removed from every dataset. The columns
first_author_emailandcorrespondence_author_emailno longer exist inwashdev,ploswater,uncnewsletter,wsorjwh. They were published up to and including v0.3.0. The addresses are personal data and earn nothing analytically, since the research questions use author country,das_type, keyword frequency and supplementary counts. A CC BY table of corresponding author addresses is a ready-made mailing list, which is the concrete harm. The addresses remain indata-raw/for provenance; code that read either column needs updating. Addresses that authors wrote into the statements themselves are masked in place, keeping the domain so the statement still reads correctly (52 statements across the five datasets).The flat-file exports in
inst/extdata/are no longer built into the installed package. They are still generated and still live in the repository and the Zenodo deposit, but shipping a CSV and an XLSX per dataset for six datasets would push the built tarball past the 5 MB CRAN guidance. Calls tosystem.file("extdata", ..., package = "washopenresearch")no longer resolve; read the datasets directly instead, for exampledata(ws).inst/CITATIONis unaffected and still ships, socitation("washopenresearch")is unchanged.Access tokens are stripped from repository URLs. A few statements carried a Zenodo pre-signed link (
?token=<JWT>) granting access to an otherwise restricted record; the token is replaced and the record URL kept, following the same reasoning as the expired Silverchair signatures in #10.
New features
Two new datasets covering the remaining IWA journals scraped in the same run as
washdev(#32-#34).wsholds all 4,884 articles of Water Supply from 2001 to 2026, andjwhholds all 2,013 articles of the Journal of Water and Health from 2003 to 2026, both collected withdata-raw/iwa_scraping.Rand sharing thewashdevschema. Together withwashdevandploswaterthis takes the corpus to 8,403 screened articles. Data availability statements were mapped to the shareddas_typelevels for 1,596 of 1,623 statements inwsand 717 of 733 injwh; the unmapped tail keeps the full statement text and is listed indata-raw/ws-das-review.csvanddata-raw/jwh-das-review.csv(#12).A fourth IWA journal, AQUA, was scraped in the same run but is not exported. 539 of its 1,819 rows carry
has_dasTRUE whiledasanddas_typeare empty, so the statement text was never captured and the dataset’s central variable would ship empty. It needs a re-scrape first.Yash Dubey is added as an author (#31). The contribution predates the R port: the Selenium-based scraper and a washdev data update, committed in December 2024 under the GitHub Action identity, which is why it was missed when the author list was last reviewed.
washopenresearch 0.3.0
New features
- New function
das_in_paper_support()classifies how a “data in paper” data availability statement is backed by the recorded supplement fields (#47). The modal claim “all relevant data are included in the paper or its supplementary information” splits into"no supplement"(the claim rests on the printed tables alone),"unstructured supplement"(pdf or images only),"structured supplement"(docx, xlsx and similar), and"open supplement"(csv, txt, json, xml). Across washdev and the three IWA journal snapshots (2,599 claims), 71.5% have no supplement, 26.5% a structured one, 2.1% an unstructured one, and none an open format.data-raw/das_in_paper_support.Rreproduces the summary indata-raw/das-in-paper-support-summary.csv. A follow-up cross-check against OpenAlex article types (data-raw/das_in_paper_article_types.R, committed DOI-to-type lookup) shows the no-supplement claims are almost entirely substantive research articles: only 0.5% are front matter or reviews, so the article-type filter proposed in #47 does not shrink the population. Verifying the remaining 1,847 bare claims requires reading the articles’ tables. - Supplement content audit (#47 tier 3): the 741 supplementary files of the IWA snapshots whose pre-signed CDN links were still valid were downloaded (checksummed manifest in
data-raw/suppfiles/manifest.csv; the signatures lapse 2026-08-18 to 2026-08-30, so 572 further links were already dead) and classified with transparent structural heuristics (data-raw/suppfiles_parse.R: pandoc-converted docx tables, readxl/xlsx sheet metrics). Of the 290 in-paper claims whose structured supplement was in hand, 54% share prose only, 21% summary tables, and 20% tables shaped like observations (data-raw/das_in_paper_suppfile_audit.R). xlsx files are the exception: 25 of 29 hold observation-shaped sheets. Heuristics are recorded per file and await validation against a manual sample. - New scripted acquisition pipeline for a fourth dataset,
datapapers, covering WASH-related data papers in seven dedicated data journals (Scientific Data, Data in Brief, Gates Open Research, F1000Research, GigaScience, GigaByte, and Data (MDPI)) (#28). The pipeline lives indata-raw/01_datapapers_acquire.R(Crossref/Europe PMC harvest with a committed raw snapshot),data-raw/02_datapapers_screen.R(relevance screening captured in a committed decision sheet keyed on DOI), anddata-raw/03_datapapers_process.R(harmonisation to the shared schema and export). The first harvest and screening round yielded 8 papers, shipped as the newdatapapersdataset with CSV and XLSX exports ininst/extdata/.
Minor improvements and fixes
-
datapapersnow carries the repository links its papers deposit to:data_repo_urlanddata_repowere NA for all 8 papers because the Crossref relation metadata was empty and the fallback planned for #27 never ran. The links were verified against Crossref relations, DataCite resource types, and the articles’ availability sections, and are recorded indata-raw/datapapers_repo_fixes.csv, applied during processing. Seven papers use general repositories (GBIF, IEEE DataPort, Figshare, Dryad, NCBI BioProject, Zenodo); none uses a WASH sector platform. - Expired pre-signed CDN links in
washdev$supp_urlare rewritten to stable DOI URLs, and Google Scholar alert redirects inuncnewsletter$paper_urlare decoded to their target URLs (#10). The 343 Silverchair links carried a January 2024 expiry, and the high-entropy signature tokens tripped secret scanners; the article DOI is recovered from the link path, so no re-collection is needed. Two helpers indata-raw/helpers.R,canonicalize_silverchair_url()anddecode_scholar_redirect(), do the rewrites reproducibly. - The list-column collapsing helper and shared country-cleaning steps moved to
data-raw/helpers.R, sourced by all processing scripts. -
data-raw/README.mddocuments the run order and provenance of every committed snapshot and decision sheet.
washopenresearch 0.2.0
New features
- New dataset
ploswaterwith all 436 articles of the journal PLOS Water from its first volume (2022) to July 2026, collected through the public PLOS API (#14). Beyond the shared schema it recordsdas_repo_url(links and dataset DOIs in the data availability statement),das_repo_name(the recognized repository behind them),article_type, andpublication_date(#15). Data availability statements are mandatory at PLOS: all 333 research articles carry one, and 177 articles link out to a data location. -
washdevnow covers volume 1 (2011) through volume 16 issue 6 (June 2026) with 1173 observations, up from 932 (#12). The 241 new rows cover volumes 14 to 16. -
washdevgains adoicolumn, filled for the newly scraped articles; earlier rows will be backfilled via Crossref (#20). - Data acquisition is now fully R. The washdev scraper was ported from Python/Selenium to
data-raw/washdev_scraping.Rusing chromote and rvest, with incremental updates (#11). The Python tooling ininst/python/, including a 17 MB chromedriver binary, was removed (#17).
Bug fixes
- The manual supplement-type corrections for
uncnewsletterwere indexed againstwashdevpaperids and landed on the wrong rows, andcorrespondence_author_affiliation_countrywas never cleaned because the cleaned values were written intofirst_author_affiliation_country, overwriting it. Both are fixed anduncnewsletterwas regenerated; country values changed on 56 rows and supplement types on 16 rows (#13). - Two author names in
washdev(for example “Inês Freire Machete”) carried Mac Roman bytes that made the xlsx export fail; the raw data is repaired at read time (#13). -
.Rbuildignoreexcluded the wholeinst/directory from the built package, socitation("washopenresearch")and theinst/extdatafiles were missing from installed packages. The rule is now scoped correctly (#17).
Minor improvements
- The data dictionary and roxygen documentation use the actual variable names
first_author_affiliation_countryandcorrespondence_author_affiliation_countryforwashdev(previously documented as*_affiliation_region) and document theurl_sourceanddoivariables. Theuncnewsletterdictionary entry forissue_urlis corrected from “Volume number of the journal” (integer) to the newsletter issue URL (character). -
uncnewsletteris documented as a frozen source: the newsletter ceased publication in May 2024 (#17). - For newly scraped
washdevrows, mixed supplementary file types are recorded as ” & “-joined lists (one type per file) instead of the literal”misc”, which previously required manual repair. - Data values that no cleaning rule could resolve are written to review files under
data-raw/(*-das-review.csv,*-country-review.csv) instead of being fixed by hand, so every correction stays in reproducible R code (#12, #15).
washopenresearch 0.1.0
Breaking changes
-
washdevanduncnewsletterno longer contain list-columns (#8). The multi-value variablessupp_file_type,supp_url,das_repo_url, andkeywordsare now character columns in which multiple values are separated by"; ". Flat-file exports such aswrite.csv()now work directly on both datasets. Code that usedtidyr::unnest()orunlist()on these columns should split the strings instead, for example withtidyr::separate_rows(supp_file_type, sep = "; ")orstringr::str_split(keywords, "; "). -
Dependswas raised from R (>= 2.10) to R (>= 3.5), required by the serialization format of the regenerated data files.
Minor improvements and fixes
- The CSV and XLSX exports in
inst/extdata/now show multi-value cells as"; "-separated strings instead of R code literals such asc("pdf", "docx"). - The variable descriptions in the package documentation and the data dictionary describe the new delimited format and the correct variable types.
- The article “Missed Opportunity: where is WASH research data gone?” uses
tidyr::separate_rows()in place oftidyr::unnest()to expandsupp_file_type. - The word cloud figure in the README has alt text.