You need to find large spatio-temporal datasets — but where?
While there are hundreds of publicly available datasets, locating them can take months of searching. When potential sources are found, they rarely provide enough information for a researcher to decide if the set actually contains the kind of data they need without downloading the often huge file and sorting through it first.
Thanks to a computer scientist at the University of California, Riverside, finding the right dataset is now as easy as bookmarking a website, and it costs absolutely nothing.
Ahmed Eldawy
Ahmed Eldawy, an assistant professor of computer science in the Marlan and Rosemary Bourns College of Engineering, and his group spent the last three years combing the internet for public spatio-temporal datasets, studying their attributes, and summarizing the results for each set on interactive maps that show the user exactly what they’re getting.
“People who work on data science need datasets but can spend a lot of time finding them,” Eldawy said. “I wanted to build an archive they can find easily.”
Called the UCR Spatio-temporal Active Repository, or UCR STAR, the archive is made available as a service to the research community to provide easy access to large spatio-temporal datasets through an interactive exploratory interface. Users can search and filter those datasets as if shopping for their research, except that everything is free.
“The map interface visualizes the data, so you can see if it’s a good fit,” Eldawy said. “It’s like a catalog for datasets.”
At the heart of UCR STAR, the map provides an interactive exploratory interface for the dataset. Similar to Google Maps or other web maps, users can zoom in and out and pan around to get a quick overview of the data distribution, coverage, and accuracy.
Important details are displayed once a dataset is selected, such as the original homepage, a link to the original download source, size in bytes, number of records, file format, and other useful information. The subset download feature allows users to quickly download the data in a given geographical region, which reduces the download size. They can also embed their customized view on a webpage or share the link via social media and bookmark it to revisit later.
UCR STAR contains 102 datasets and 5 billion records. The datasets were mapped using Da Vinci, an open source framework built on top of Apache Spark that Eldawy designed to work with spatial data. The UCR STAR website is best accessed through a desktop browser but also has a limited mobile-friendly interface.
Gary Price (gprice@gmail.com) is a librarian, writer, consultant, and frequent conference speaker based in the Washington D.C. metro area.
He earned his MLIS degree from Wayne State University in Detroit.
Price has won several awards including the SLA Innovations in Technology Award and Alumnus of the Year from the Wayne St. University Library and Information Science Program. From 2006-2009 he was Director of Online Information Services at Ask.com. Gary is also the co-founder of infoDJ an innovation research consultancy supporting corporate product and business model teams with just-in-time fact and insight finding.
From the Associated Press: A roundup of some of the most popular but completely untrue stories and visuals of the week. None of these are legit, even though they were ...
From The Sydney Morning Herald: Authors, illustrators, and editors will be compensated for e-book and audiobook library borrowings for the first time, in a move by the federal government to ...
From the National Archives and Records Administration (NARA): A draft Customer Research Agenda was open for public review and comment in October 2022. “We’re grateful for the feedback we received ...
From MIT Technology Review: Hidden patterns purposely buried in AI-generated texts could help identify them as such, allowing us to tell whether the words we’re reading are written by a ...
From the Congressional Research Service: Nearly one in four Americans has a disability, according to 2018 estimates from the U.S. Census Bureau. Congress has recognized that in addition to making ...
From The NY Times: When [Joan] Didion died in 2021 at age 87, the news set off an outpouring of tributes to a writer who fused penetrating insight and idiosyncratic personal voice, ...
Below, Find the Full Text of a Letter Sent to the Carolina Community From Kevin M. Guskiewicz University of North Carolina at Chapel Hill Chancellor Kevin M. Guskiewicz and J. ...
From the Boston Public Library: The Boston Public Library is proud to contribute to the celebration of Black History Month with its annual “Black Is…” booklist. The booklist aims to commemorate ...
From NYU Langone: Researchers at NYU Grossman School of Medicine, in partnership with the Robert Wood Johnson Foundation, unveiled the Congressional District Health Dashboard (CDHD), a new online tool that ...
From a cOAlition S Announcement: Transformative arrangements – including Transformative Agreements and Transformative Journals – were developed to encourage subscription journals to transition to full and immediate open access within a defined timeframe (31st December 2024, ...
From the Library of Congress: The Library of Congress announced today the appointment of Hannah Sommers as the new Associate Librarian for Researcher and Collections Services in the Library Collections and Services Group. In this role, Sommers will lead the future of the Library’s collections and the services it delivers to researchers and users. She will be central ...
As Book Bans Increase Across the Country, a Boston University Scholar is Fighting Back Core’s Library Resources & Technical Services Journal Goes Fully Open Access Digital Image Processing: It’s All ...