Tag Archives: Palladio

Reflections: DPHF17 WHCC Mapping 1.0

I. Aim of the Project

The aim of the DPHF17 WHCC Mapping 1.0 project is to apply what I am learning in the George Mason University Digital Public Humanities Certificate program to the larger historical project “Women and Health Care in Cleveland” (WHCC).  WHCC has examined many types of ‘connections’ between Cleveland’s institutions and individuals using traditional historical methods and disseminating this information through academic and public avenues.  As the WHCC project evolved, it has expanded including many variables: Public and private institutions, corporate, proprietary and government entities, relationships that span the gamut of intimate families and friends, through organizational and philanthropic to distant resource and approval figureheads.  The increasing size of the corpus data has made digital humanities methods attractive to the WHCC.

For this specific project, I am applying two forms of mapping and network analysis software to a small sample WHCC data set to ascertain our ability to adapt WHCC materials to DH methods. I am anxious to see what these visualizations reveal about the WHCC network of medical care in Cleveland, OH during the period 1850-1870, when the city and public institutions are just organizing.  This project is also being conducted as part of a learning community created within the Digital Humanities Certificate Program.  Collaboration is an intrinsic part of projects created within the field of digital humanities. Close engagement with persons involved in all aspects of the resourcing, researching, production and those who will be engaging with the completed research is part of this methodology. (Ramsay, Stephen, “DH Types”) The production of this project and its partners is an exercise in collaboration.  Also this effort will inform WHCC current and future research and project activities.

II. Sources

WHCC records are comprised of traditional research notes, tabulated and compiled over years of assorted projects and presentations.  These records needed to be transformed into complete, structured and clean databases to form a Subset of data for digital analysis.  The technical process of converting WHCC information from its original Excel Spreadsheet to a .cvs database ready to be utilized for DH analysis is described on the WHCC Project Page.

For this particular reflection, it needs to be recognized how “incomplete” the corpus was from a digital standpoint.  This is part of the “art” of traditional historical analysis – interpreting relationships from ‘limited’ sources.  Various specific details are missing for a majority our people, but the ‘missing’ data is not always the same aspect.  The lack of quantifiable and definitive data results in the critiques, such as “Most historians still trade in text mediated by uncertainty and theory; while historical geographers, strive to tie data to a knowable and certain fragment of the world’s surface”, by social scientists that have more measureable data on which to rest their authority. (Hitchcock, Historyonics)  The time consuming task of transforming the minutia of notes (fortunately frequently recorded in digital formats – Word, Excel, Google etc.) was tedious and required several passes to add  missing elements,  correct items, and standardize references. The purpose as Hitchcock asserts

“makes possible a situation in which historians cease to be mere text merchants, obsessed with the perfect quote, and compelling (if largely un-evidenced) argument; and where geographers {AND HISTORIANS} have a new access to the subtle mappings of the marks of ‘culture’ in its broadest sense – a new way of thinking about the geographical distribution of behaviors and ideas, that bring within a geographical fold questions traditionally preserved for others…..We quite suddenly share a new culture of data – and data can be translated.”  (Hitchcock, Historyonics)

But, recorded in these records were extensive descriptions of relationships, references and speculations about connections that cannot be translated into spreadsheet columns.  This difficulty has been noted:

“Humanistic data are almost by definition uncertain, open to interpretation, flexible, and not easily definable. Node types are concrete; your object either is or is not a book. Every book-type thing shares certain unchanging characteristics.  … This reduction of data comes at a price, one that some argue traditionally divided the humanities and social sciences. If humanists care more about the differences than the regularities, more about what makes an object unique rather than what makes it similar, that is, the very information they are likely to lose by defining their objects as nodes.”…..

“Unfortunately, given that humanistic data are often uncertain and biased to begin with, every arbitrary act of data-cutting has the potential to add further uncertainty and bias to a point where the network no longer provides meaningful results. The ability to cut away just enough data to make the network manageable, but not enough to lose information, is as much an art as it is a science.” (Weingart, “Demystifying Networks”).

An objective of this project is to create a physical picture, while maintaining our awareness that this “picture” is a created artifact.  The primary criteria of removal from our corpus was incomplete addresses, this reduced the initial data set from 772 to 194.  Identifiable geolocations reduced it to 135.  In statistical terms, the test data is comprised of 17% of the full corpus.  The analysis has also been governed more by ‘quantifiable’ units (nodes) and is therefore more broad and superficial than the depth of WHCC records indicate.  For example, Oliver Hazard Payne, the grandson of Oliver Hazard Perry, served in the Civil War and then became a board member of Standard Oil.  He donated to Cornell Medical College, Cleveland’s Lakeside Hospital, the Jewish Orphan Asylum, St. Vincent Charity hospital among others, but died a bachelor.  He left his money to family members including Elizabeth Bingham Blossom, who married Dudley Blossom, who would become the City Welfare Director and gained national prominence transforming “City” into Metropolitan Hospital.  — How does that get put in a spreadsheet?

“There are many networks that can connect the same group of nodes, and many questions that can be asked of any given network, but before trying to use networks to study history, you should be careful to make sure the questions match the network.” (Weingart, Networks Demystified).  In order to accomplish this test within the limited time-frame, specific decisions were made to include rows that had complete data and columns that could be straightforwardly quantified or compared digitally.  These choices dramatically limited the size of the Trial Data Set and the number of visualizations that could be performed.  It is evident that experience in database design would enhance the structure of the database and the end-product.

III. Software that You Worked With:

The tools used for analysis in the DPHF17 WHCC Mapping Project were CARTO and PALLADIO.  A description of the tool and examples of the visualizations produced with them is available on the WHCC Project Page .

There are aspects of “working” in the field of Digital Humanities that completely change your conception of it.  I recognize “plural” in “Digital Humanities” in a much different way than when I began the semester.  At that point, I believed it was interdisciplinary and combined the digital (discrete bits & computers) with humanistic topics.  The primary purpose being to create products and exhibits that contributed to the understanding and snazz of the Humanities. The Digital Humanities are plural in a much deeper and comprehensive way.  In the creation of the DPHF17 WHCC Mapping Project, I have used multiple softwares.  I think a more appropriate way of discussing them is in terms of categories of tools used, as well as specific applications.

Throughout the project process, communication of ideas and transmission of data has been necessary and on-going.  However a straight forward email is only one form.  Daily I was using several platforms based on where I was working and the type of collaboration I was seeking: Microsoft Outlook at work; Gmail for private and college business communications; Blackboard messages for student interaction within my class, as well as blackboard collaboration for class meetings, Slack – was the primary collaboration tool used for course questions and help within our group.  Each one of these platforms is directed to accomplishing similar goals but in targeted and specific ways.  Blackboard (An education management system) is a suite of tools to address the needs of educators and students.  Microsoft Outlook is particularly targeted to business processes and Gmail tries to straddle both business and private communication spheres.  Choice of usage is not always in the hands of the user, but may be dictated by the environment.

Collection and organization of data was also accomplished through many systems.  The course itself had its own course platform “dhcert” which was customized software hosted on Reclaim.  This enabled the dispersal of information and submission of coursework in a paced process that was much more fluid and individually accessible than Blackboard.  I utilized document tools –  Google Docs and Microsoft Word; spreadsheet tools – Excel and Google sheets.  The primary advantage of Google tools being the ability to synchronously edit and collaborate.

Tools for the dissemination of results include: WordPress – This is publishing software which is hosting my blog and also provides a course record.  The capacity in which it was used this semester was more an educational and professional platform than personal journal.  Also discussed and encouraged are further social media tools utilized to “get the word out”, now that newspapers and print media are not as commonly utilized.  These tools included Facebook, Twitter, other people’s blogs, institutional websites and professional networking resources like Linked In and HNet (field specific network) or sites like www.digitalhumanities.cam.ac.uk

The WHCC Project is a ‘location specific’ project.  At various times, I have hand created maps or adapted existing maps to “show” the relationships with which we are working; and how space impacts those relationships.  So I was drawn to the mapping and visualization tools of CARTO and Palladio.  Multiple tools were selected, because I am interested in testing both the visual map interface and the network constellation visual representations.  Texas A&M GeoServices was an essential tool as it was used to transform street addresses into latitude/longitude coordinates as Geo locations. This opensource program was particularly friendly to work with, utilized my Excel sheet with expanded columns formulated the Geo locations for my limited data set quickly.  The Excel sheet was saved as a .cvs file, which once created could be utilized by both platforms, which was a pleasant surprise.  CARTO maps.  It is possible to upload the GIS data, have it appear on a constructed map.  CARTO then allows the user to make various forms of analysis (Histograms, Graphs, Time Series, Density and Grouping) that help show connections and changes over time on the map itself.  CARTO is easy in the sense that its widgets and analysis software will do magic in transforming data into maps.  But it is a time-consuming process to learn CARTO’s idiosyncrasies and getting the various widgets to operate the way you think they should intuitively. CARTO I can utilize 19th century maps digitized in the Cleveland Map collection. I am choosing this software because, we will be looking at this network again in the 1930’s and 1990’s and utilizing the GIS aspects of CARTO will be particularly helpful.

The Palladio platform is easier to work with than CARTO.  However, the ability to customize visualizations is more limited.   Palladio is structured more as a “tool set”.  The tool is selected, the data is plugged in and “poof” the network image appears.  This makes it much more fun for the non-techie to do techie magic.  In projecting a new form of visualization which focuses on the connection and its strength; Palladio helps the user focus on the relationship between objects.  This non-traditional spatial view of relationships helps the user to consider the meaning of the patterns and stimulates new questions.   Using Palladio and CARTO together

 IV.  Problems Completing the Project:

Overall the major challenge of the project was time.  The time-frame of this sub-project and many design elements have been dictated by DPHCert course requirements.  A substantial portion of time for this project has been spent reorganizing, cleaning and supplementing the information that had been gathered in order to create a clean data set that could be loaded into visualization software.  The scope of the current project has been narrowed from a full scale analysis of the WHCC data to testing a portion of the data. Even then, the scope was reduced further as each activity in the project required learning the methodology of a new tool, or expanding on a fragmentary knowledge of it.  While time available to the project was intended to focus on utilization of the digital tools, the preservation of the larger project’s resources became a time critical issue. While this aspect seems extraneous to the DPHF17 project, it is a valuable reminder of human, machine and societal issues that will impact any public project.  Another ‘human’ issue in the completion of this project is isolation from an active DH facility/community.  As I read fellow “ground” students posts, which included references to conversations and evidently lab type resources and assistance, that resulted in professional products; I recognized the depth of my struggle to learn unfamiliar tools and change embedded practices on my own.

V.  Discoveries About Sources:

Preparation of the full corpus data was beyond my ability to complete in the time available in the semester. The Test corpus was determined by a representative sample and having a substantial number of records to process. Limiting the number of columns in the network relationships (determined as necessary during the data cleaning process) was the most disappointing feature of creating the test data. Also after having compiled the test corpus and run the visualizations, I believe I will need to rethink strategy of formulation. Names are important to understanding the relationships in the larger project, but the ‘significance of the names’ is invisible to the computer. The addresses are also significant, but do not tell the full story of the geographic locations.

WHCC spreadsheets were created to compile resources for the purposes of other  sub-projects. This format was not sufficient for effective digital processing. The lack of a fundamental design structure meant that creating a “usable” database was more complex and time consuming than expected. Gaps in the data needed to be filled, extra punctuation and side notes needed to be removed. Addresses has to be placed into standard formats and geolocations calculated. Computer tools can be used for some of these processes, but the interpretation of meaning (even in deciphering aspects within the data structure) requires human analysis.

Choices were made to limit the scope of the project because data was missing or incomplete, but I also discovered that some of the organizational structure even in the newly constructed database was unhelpful.  The original construction of the numbered Charity columns was based on the evidence of commitment or office tenure of the individual.  For the purposes of digital analysis, I believe the columns should be titled by Charity Name.  This would allow a better display of the charities as well as the individual relationships.  Additional dates, (hiring, membership, terms of office etc.) if available would make time sequence analysis possible.

Opening the project up to the public during its ‘working phase’ is not a strategy I would have anticipated or probably considered. The inclusion of public awareness is part of the evolution of my understanding of ‘Digital Humanities’ and exposure to more DH work in a thoughtful way. At our first glance of Digital Harlem, I realized the power and usefulness of its blog. Since this project is somewhat similar to WHCC, in exploring the impact of an aspect of an urban community; I examined this project closely and have been reflecting on its strategy and methodology. WHCC does have a ‘crowdsourcing’ element in its origin and full implementation. Laying the groundwork for that engagement now, seems appropriate,

“[DH] projects allow improved searching and analysis of primary sources, facilitating a researcher in their task and allowing new ways to synthesize, juxtapose, and create knowledge. Although many digitization projects in the Digital Humanities are developed with a core constituency in mind, the resulting resources can be used by a wider audience, and the ramification of opening up access to humanities, arts, and cultural content is just starting to be understood.” (Terras, Digitization and Digital Resources in the Humanities)

 VI.  Highlights of Presentation

WHCC has been collaborative from the beginning in the research, generating community interest and support.  This project has benefitted from and reinforced the significance of that aspect of WHCC practice.  WHCC is an effort to uncover the networks and contributions of previously unrecognized actors in the development of our current health care system.  This project has demonstrated the ability of Digital Humanities methodology and tools to uncover new relationships and generate import research questions. Production of Digital Humanities analysis is time-consuming (expensive) on the front end, but replicable on the back end.

“It is not just that mapped data is seen in its geographical context…large quantities of data, can be combined on a single map, providing an image of the complexity of the past.  You can examine maps of sources at different scales, and discover relationships…by visually detecting spatial patterns that remain hidden in texts and tables.” (Robertson, “Digital Harlem”)

The introduction of this methodology into WHCC proposes a broader understanding of the ‘unrecognized actors’ and a deeper understanding of the embedded nature of the networks within the city.  “Trevor Harris, Jesse Rouse and Susan Bergeron argue, ‘The visual display of information creates a visceral connection to the content that goes beyond what is possible through traditional text documents.’”(In Robertson, “Digital Harlem”)  This project is a step in developing a ‘visceral connection’ within the WHCC.

Notes:

Edelstein, Dan, Findlen, Paula, Ceserani, Giovanna, Winterer, Caroline, and Nicole Coleman. “Historical Research in “A Digital Age: Reflections from the Mapping the Republic of Letters Project.” American History Review, 122, no. 2 (2017): 400-424.

Robertson, Stephen “Putting Harlem on the Map.” In Writing History for the Digital Age.  Jack Dougherty and Kristen Nawrotzki, ed.  University of Michigan Press, 2012.

Hitchcock, Tim. “Place and the Politics of the Past.” Historyonics (blog), July 11, 2012.

Ramsay, Stephen. “DH Types One and Two”. Stephen Ramsay (Blog), 2013.

Terras, Melissa.  “Digitization and Digital Resources in the Humanities.” In Digital Humanities in Practice, ed. Claire Warwick, Melissa Terras, and Julianne Nyhan.  Facet Publishing, 2012.

Weingart, Scott. “Demystifying Networks” Scott Weingart (blog), December 14, 2011.

Weingart, Scott. “Networks Demystified 8: When Networks are Inappropriate.” Scott Weingart (blog), December 5, 2013.

Playing with Palladio

Just like a kid at a table, I was given some preformed clay and told to make some beautiful works.  Not having to do “the set up”, Palladio was relatively easy to get “into” network analysis and start messing around.   Palladio is a “general-purpose suite of visualization and analytical tools” that enables researchers to map and graph humanities data. (http://hdlab.stanford.edu/palladio/about/)  When we are discussing data here we mean “corpus” (techie term – compilations for the rest of us) of, in my case, the historian’s research (people, events, addresses, dates, contacts, documents, etc.).   The data can be vast amounts of material (text, photographs, tables, etc.) or relatively small collections.   The software in the platform through the use of algorithms (sets of directions) transforms our element (node) into a “plotable point on a map or graph.  The schema for assembling, disassembling and reassembling produces various kinds of maps, graphs, constellations all with the associated metadata accessible.  These manipulations draw (“edges” for the techies) “connections and patterns” (for the rest of us) that can be seen.   These “visualizations” foster enhanced understanding of the data itself and make obvious how elements within the data are connected or not and because of the ability to “play” (isolating “ALL” the variables and being able to test relationships or time series, in space or against a physical map) new questions are brought to the researchers mind, hypothesis can be verified and various problems or possibilities can be explored and addressed.  Our ‘Ideas’ become pictures that enable us to work with them, without losing track of them, from various perspectives.

The designers of Palladio built it for mapping the correspondence metadata of a project called “The Republic of Letters” which explores the connections of Enlightenment Intellectuals and  the transference of ideas through their letters.  (Visualize massive amounts of text in multiple languages).   Palladio understands the historian’s mind and methodology,

“It is an environment that supports thinking through data. Like so many historical sources, the (data) is incomplete in unpredictable ways. Historians work with fragments of information from the past, not complete data sets. ..tools need to support scholars in building an understanding of the historical material through working with it. To create Palladio, we had to externalize what is a very internal, individual, thought process. “  http://hdlab.stanford.edu/palladio/about/

My trial experience working with Palladio centered around a data set created by Prof. Robertson from the Works Progress Administration Slave Narratives. (https://www.loc.gov/collections/slave-narratives-from-the-federal-writers-project-1936-to-1938/about-this-collection/ ) Since I was able to upload ‘Palladio prepared data’, I was able to avoid a time consuming task of aligning my data for Palladio manipulation.

It was relatively fast and easy to create the primary table (main data set) against which other tables (elements of the data) could be manipulated.  For the Slave Data, the primary data set was the Alabama Slave Interviews and the subtables were identifications of the former slaves that were interviewed and data about the interviewers.  Once the tables are created, they can be analyzed with various tools including locating the slaves/interviews on a map, drawing connections between an interviewer and the slaves interviewed, creating network analysis of various types by the individual elements within the data sets.  The “visualizations” I made included maps, string and cloud diagrams (okay I admit– I got interested and played with time series and graphs as well).   It was fascinating and time consuming to try different arrangements that suggested unique relationships, impact of place and Slave occupation.  Also revealed were limits in the veracity of the data (not a reliable reflection of slave tasks or a true reflection of statewide representation), but also new issues to consider, such as how does the gender of the interviewer impact the questions used in an oral history interview?

Palladio describes itself as “a tool for reflective practice.”  Both Scott Weingart (“Networks Demystified” Blog December 5, 2013) and Ryan Cordell (“Reprinting, Circulation and the Network Author in Antebellum Newspapers”),  call to our attention that compiled data has its issues, in the same way normal archival research has gaps and biases and snags.  It is easy to see how visualizations could cause erroneous interpretations.  And truly the data is just a picture of the facts and still requires interpretation as to its meaning and significance.  So the mind and integrity of the historian is still crucial to good history…but new Playdough is always fun and creative!