Tag Archives: Text Analysis

Playing with Palladio

Just like a kid at a table, I was given some preformed clay and told to make some beautiful works.  Not having to do “the set up”, Palladio was relatively easy to get “into” network analysis and start messing around.   Palladio is a “general-purpose suite of visualization and analytical tools” that enables researchers to map and graph humanities data. (http://hdlab.stanford.edu/palladio/about/)  When we are discussing data here we mean “corpus” (techie term – compilations for the rest of us) of, in my case, the historian’s research (people, events, addresses, dates, contacts, documents, etc.).   The data can be vast amounts of material (text, photographs, tables, etc.) or relatively small collections.   The software in the platform through the use of algorithms (sets of directions) transforms our element (node) into a “plotable point on a map or graph.  The schema for assembling, disassembling and reassembling produces various kinds of maps, graphs, constellations all with the associated metadata accessible.  These manipulations draw (“edges” for the techies) “connections and patterns” (for the rest of us) that can be seen.   These “visualizations” foster enhanced understanding of the data itself and make obvious how elements within the data are connected or not and because of the ability to “play” (isolating “ALL” the variables and being able to test relationships or time series, in space or against a physical map) new questions are brought to the researchers mind, hypothesis can be verified and various problems or possibilities can be explored and addressed.  Our ‘Ideas’ become pictures that enable us to work with them, without losing track of them, from various perspectives.

The designers of Palladio built it for mapping the correspondence metadata of a project called “The Republic of Letters” which explores the connections of Enlightenment Intellectuals and  the transference of ideas through their letters.  (Visualize massive amounts of text in multiple languages).   Palladio understands the historian’s mind and methodology,

“It is an environment that supports thinking through data. Like so many historical sources, the (data) is incomplete in unpredictable ways. Historians work with fragments of information from the past, not complete data sets. ..tools need to support scholars in building an understanding of the historical material through working with it. To create Palladio, we had to externalize what is a very internal, individual, thought process. “  http://hdlab.stanford.edu/palladio/about/

My trial experience working with Palladio centered around a data set created by Prof. Robertson from the Works Progress Administration Slave Narratives. (https://www.loc.gov/collections/slave-narratives-from-the-federal-writers-project-1936-to-1938/about-this-collection/ ) Since I was able to upload ‘Palladio prepared data’, I was able to avoid a time consuming task of aligning my data for Palladio manipulation.

It was relatively fast and easy to create the primary table (main data set) against which other tables (elements of the data) could be manipulated.  For the Slave Data, the primary data set was the Alabama Slave Interviews and the subtables were identifications of the former slaves that were interviewed and data about the interviewers.  Once the tables are created, they can be analyzed with various tools including locating the slaves/interviews on a map, drawing connections between an interviewer and the slaves interviewed, creating network analysis of various types by the individual elements within the data sets.  The “visualizations” I made included maps, string and cloud diagrams (okay I admit– I got interested and played with time series and graphs as well).   It was fascinating and time consuming to try different arrangements that suggested unique relationships, impact of place and Slave occupation.  Also revealed were limits in the veracity of the data (not a reliable reflection of slave tasks or a true reflection of statewide representation), but also new issues to consider, such as how does the gender of the interviewer impact the questions used in an oral history interview?

Palladio describes itself as “a tool for reflective practice.”  Both Scott Weingart (“Networks Demystified” Blog December 5, 2013) and Ryan Cordell (“Reprinting, Circulation and the Network Author in Antebellum Newspapers”),  call to our attention that compiled data has its issues, in the same way normal archival research has gaps and biases and snags.  It is easy to see how visualizations could cause erroneous interpretations.  And truly the data is just a picture of the facts and still requires interpretation as to its meaning and significance.  So the mind and integrity of the historian is still crucial to good history…but new Playdough is always fun and creative!

“Seeing” the Slave Narratives with Voyant

Voyant Tools  is a web-based text reading and analysis environment designed to facilitate reading and interpretive practices for digital humanities and the public.  http://voyant-tools.org/docs/#!/guide/about

Voyant Tools is an entry point – software that allows you to do basic computer-assisted analysis of data.

Text or edited and prepared data can by studied and analyzed in new ways and ‘plotted’ so that it can be saved and shared with others.   Using Voyant allows others to explore your data through interactive panels that can be inserted into online publications and presentations.   Voyant can also be used as a development platform to design new forms of analysis and visualization.  The main visualizations are: word frequencies, word clouds and “distinctive words”.

I am also including a link to the scatterplot for this corpus.  Clicking on the “Analysis” option > Principal Component Analysis provides  further evidence of the significance of “race” as a feature of slave experience.   http://voyant-tools.org/?corpus=2d0c91a1da599627fea6d35d607795a0&view=scatterplot&stopList=keywords-96814a63d28a5e281f75e02c891dcdbe


The trends graph indicates the significance of location.

I found Voyant to be an interesting tool which will be of great use when texts are clearly comparative.  It is great for collating data!  It provides the ability to analyze text to see trends in language and usage.  The word cloud is interesting and illustrative.  In the Slave Narrative data set on which I was working the common words were Old (which I ignored as the subjects of the text were elderly), Come and Time.  Whereas I would have expected more direct slave references which were secondarily popular.  The nature of the documents (by state) also provided insight into regional variations, but also made me recognize the states from which the largest number of narratives was drawn, would bias the representations and conclusions.  Understanding the Data (Musher, The Other Slave Narratives) is critical to recognizing the significance of patterns.  Having a larger context besides the quantitative data is essential for interpretation.  Also it became evident as I was working with the Data, that my choices were subjective and these choices would also “skew” the results…my lack of technical prowess also would play a role in the accuracy or distortion of the results.

For the new user….agh!  I am sure this tool would be easy for humanists with digital strengths.  For the digitally challenged, it is probably pablum, and a good entry point to Digital Mining and Modeling tools…but we are entering a “Brave New World” where the faint hearted must learn to persevere.  I believe spending some time in the on-line guides would be very helpful.  I played around the examples and information on Voyant’s website, but before attempting a larger project – I would try to learn to manage and navigate the tool better.  Time being an issue, I could not do that.  My simple tasks became time consuming, because if I tried to backtrack or change some aspect of the data my progress was lost.  Trying to return to old screens or navigate between the tools was not as “friendly” as I hoped.  Transforming from a historian who uses traditional sources of evidence to build up thesis and wider connections – I miss the library shelf (Underwood, Theorizing Research Practices).   it is a dramatic change to reverse and drill down into the minutia of checked boxes, single characters and itemizing the search.   The good new is that when you get interesting graphs or can make unexpected connections and conclusions…ya feel great!

“Learning is not attained by chance it, it must be sought for with ardor and diligence.” Abigail Adams