SUBSCRIBE
SUBSCRIBE
EXPLORE +
  • About infoDOCKET
  • Academic Libraries on LJ
  • Research on LJ
  • News on LJ
  • Skip to primary navigation
  • Skip to main content
  • Skip to primary sidebar
  • Libraries
    • Academic Libraries
    • Government Libraries
    • National Libraries
    • Public Libraries
  • Companies (Publishers/Vendors)
    • EBSCO
    • Elsevier
    • Ex Libris
    • Frontiers
    • Gale
    • PLOS
    • Scholastic
  • New Resources
    • Dashboards
    • Data Files
    • Digital Collections
    • Digital Preservation
    • Interactive Tools
    • Maps
    • Other
    • Podcasts
    • Productivity
  • New Research
    • Conference Presentations
    • Journal Articles
    • Lecture
    • New Issue
    • Reports
  • Topics
    • Archives & Special Collections
    • Associations & Organizations
    • Awards
    • Funding
    • Interviews
    • Jobs
    • Management & Leadership
    • News
    • Patrons & Users
    • Preservation
    • Profiles
    • Publishing
    • Roundup
    • Scholarly Communications
      • Open Access

January 12, 2015 by Gary Price

University of North Carolina Library Launches DocSouth Data, Historic Texts Ready For Analysis

January 12, 2015 by Gary Price

Cool from Chapel Hill. Kudos to the UNC Library on providing this new service. Superb idea.
From the UNC Library Blog:

The newly-released DocSouth Data makes the full text of hundreds of nineteenth-century books and pamphlets available for easy download as text-only files. The materials come from four text-heavy Documenting the American South collections: The Church in the Southern Black Community; First-Person Narratives of the American South; Library of Southern Literature; and North American Slave Narratives.
“Researchers who want to experiment with text analysis have plenty of tools to choose from, but they often hit a roadblock when it comes to finding collections that are ready to be analyzed,” says Stewart Varner, UNC’s digital scholarship librarian.
[Clip]
Varner and his colleagues at the UNC Library realized that DocSouth provides a rich source of ready-to-use full-text that is of high interest and high quality.
The North American Slave Narratives collection, for example, includes every known autobiographical narrative of fugitive or former slaves published in English up to 1920. Its titles are among the most frequently used in DocSouth. The collection offers a wide-ranging picture of the experiences of former slaves and African American life in antebellum America.
Moreover, these files were transcribed by hand for the original digital files. That means they are unusually accurate, free of the errors that optical character recognition (OCR) often produces.

Collections Available via DocSouth Data

  • The Church in the Southern Black Community
  • First-Person Narratives of the American South
  • Library of Southern Literature
  • North American Slave Narratives

From the DocSouth Data Web Page

Doc South Data is an extension of this original goal and has been designed for researchers who want to use emerging technology to look for patterns across entire texts or compare patterns found in multiple texts. We have made it easy to use tools such as Voyant to conduct simple word counts and frequency visualizations (such as word clouds) or to use other tools to perform more complex processes such as topic modeling, named-entity recognition or sentiment analysis.

Here’s What You’ll Find in a Data File (Once You Download and Unzip):

  • A folder containing each of the texts in the collection as plain text;
  • A folder containing each of the texts with complete TEI/XML markup;
  • A Table of Contents (TOC) file which is a .csv file (which can be opened in Excel as a spread sheet);
  • A file neamed “Read Me;”
  • A file named “text-only.xsl.”

The plain text files can be used in text mining projects such as topic modeling, sentiment analysis and natural language processing.
The TEI/XML files have been included for advanced users who would like to isolate particular parts of text for analysis.
The .csv file acts as a table of contents for the collection and includes Title, Author, Publication Date and webaddresses for both the original Doc South page and a unique url pointing to a web accessible version of the text (this is particularly useful for use with Voyant).
The Read Me file provides documentation about the collection including information about and quirks or idiosyncrasies researchers will need to be aware of when they are working with the data.
The text-only.xsl file is the script that was used to create the folder.

Read the Complete Blog Post (with Text Analysis Graphic)
Direct to DocSouth Data Info and Downloads
Direct to the DocSouth Web Site (Browse, Search the Collections)

Filed under: Data Files, Libraries, News, Patrons and Users

SHARE:

About Gary Price

Gary Price (gprice@gmail.com) is a librarian, writer, consultant, and frequent conference speaker based in the Washington D.C. metro area. He earned his MLIS degree from Wayne State University in Detroit. Price has won several awards including the SLA Innovations in Technology Award and Alumnus of the Year from the Wayne St. University Library and Information Science Program. From 2006-2009 he was Director of Online Information Services at Ask.com. Gary is also the co-founder of infoDJ an innovation research consultancy supporting corporate product and business model teams with just-in-time fact and insight finding.

ADVERTISEMENT

Archives

Job Zone

ADVERTISEMENT

Related Infodocket Posts

Report: "Australian Authors to Receive Compensation for E-Book Loans for First Time"

From The Sydney Morning Herald: Authors, illustrators, and editors will be compensated for e-book and audiobook library borrowings for the first time, in a move by the federal government to ...

National Archives and Records Administration (NARA) Publishes  Customer Research Agenda

From the National Archives and Records Administration (NARA): The National Archives and Records Administration (NARA) has posted its . A draft Customer Research Agenda was open for public review and comment ...

Report: "A Watermark for Chatbots Can Expose Text Written by an AI"

From MIT Technology Review: Hidden patterns purposely buried in AI-generated texts could help identify them as such, allowing us to tell whether the words we’re reading are written by a ...

The Accessibility of Federal Information and Data: A Brief Overview of Section 508 of the Rehabilitation Act (Updated...

From the Congressional Research Service: Nearly one in four Americans has a disability, according to 2018 estimates from the U.S. Census Bureau. Congress has recognized that in addition to making ...

NY Times: "New York Public Library Acquires Joan Didion’s Papers"

From The NY Times: When [Joan] Didion died in 2021 at age 87, the news set off an outpouring of tributes to a writer who fused penetrating insight and idiosyncratic personal voice, ...

University of North Carolina at Chapel Hill: María Estorino Named Vice Provost for University Libraries and University Librarian

Below, Find the Full Text of a Letter Sent to the Carolina Community From Kevin M. Guskiewicz University of North Carolina at Chapel Hill Chancellor Kevin M. Guskiewicz and J. ...

Boston Public Library Celebrates Black History Month with Annual “Black Is…” Booklist & Special Events

From the Boston Public Library: The Boston Public Library is proud to contribute to the celebration of Black History Month with its annual “Black Is…” booklist. The booklist aims to commemorate ...

Research Resources: New Online Tool Provides Health Snapshot of All 435 U.S. Congressional Districts (Congressional District Health Dashboard)

From NYU Langone: Researchers at NYU Grossman School of Medicine, in partnership with the Robert Wood Johnson Foundation, unveiled the Congressional District Health Dashboard (CDHD), a new online tool that ...

Report: "cOAlition S Confirms the End of Its Financial Support for Open Access Publishing Under Transformative Arrangements After...

From a cOAlition S  Announcement: Transformative arrangements – including Transformative Agreements and Transformative Journals – were developed to encourage subscription journals to transition to full and immediate open access within a defined timeframe (31st December 2024, ...

Library of Congress: Hannah Sommers Appointed New Associate Librarian for Researcher and Collections Services

From the Library of Congress: The Library of Congress announced today the appointment of Hannah Sommers as the new Associate Librarian for Researcher and Collections Services in the Library Collections and Services Group. In this role, Sommers will lead the future of the Library’s collections and the services it delivers to researchers and users. She will be central ...

Virginia Tech: University Libraries Dean Tyler Walters Appointed Board Chair of Academic Preservation Trust; IEEE Computer Society 2023...

As Book Bans Increase Across the Country, a Boston University Scholar is Fighting Back Core’s Library Resources & Technical Services Journal Goes Fully Open Access Digital Image Processing: It’s All ...

Funding: Library Freedom Project Receives $1 Million Grant Award From the Mellon Foundation to Advance Critical Privacy and...

Here’s the Full Text of the Library Freedom Project (LFP) Announcement:   Library Freedom Project (LFP) has been awarded $1,000,000 from the Mellon Foundation to expand the program’s work. For ...

ADVERTISEMENT

FOLLOW US ON TWITTER

Tweets by infoDOCKET

ADVERTISEMENT

This coverage is free for all visitors. Your support makes this possible.

This coverage is free for all visitors. Your support makes this possible.

Primary Sidebar

  • News
  • Reviews+
  • Technology
  • Programs+
  • Design
  • Leadership
  • People
  • COVID-19
  • Advocacy
  • Opinion
  • INFOdocket
  • Job Zone

Reviews+

  • Booklists
  • Prepub Alert
  • Book Pulse
  • Media
  • Readers' Advisory
  • Self-Published Books
  • Review Submissions
  • Review for LJ

Awards

  • Library of the Year
  • Librarian of the Year
  • Movers & Shakers 2022
  • Paralibrarian of the Year
  • Best Small Library
  • Marketer of the Year
  • All Awards Guidelines
  • Community Impact Prize

Resources

  • LJ Index/Star Libraries
  • Research
  • White Papers / Case Studies

Events & PD

  • Online Courses
  • In-Person Events
  • Virtual Events
  • Webcasts
  • About Us
  • Contact Us
  • Advertise
  • Subscribe
  • Media Inquiries
  • Newsletter Sign Up
  • Submit Features/News
  • Data Privacy
  • Terms of Use
  • Terms of Sale
  • FAQs
  • Careers at MSI


© 2023 Library Journal. All rights reserved.


© 2022 Library Journal. All rights reserved.