Using Textexture to Map the State of the Union Address
This tangle of differently-sized nodes and connections is actually all the meaningful words from President Obama's latest State of the Union Address, mapped according to their associations to one another. This visualization was created using the tool Textexture, a javascript-based program that scans the full text of anything—speech, book, news article, etc.—and extracts the words that have to do with the subject matter. The tool's developer, Nodus Labs, has this to say about why the tool was created:
When we read a text, we normally follow it in quite a linear fashion: from left to right, from top to bottom. Even when we skim articles quickly online, the trajectory is still the same. However, this is not the most efficient method of reading: in the age of hypertext we tend to create our own narratives using the bits and pieces from different sources. This is an easy task with short Tweets or Facebook posts, but it becomes much more difficult when we’re dealing with newspaper articles, books, scientific papers. The amount of information we’re exposed to increases from day to day, so there’s a challenge of finding the new tools, which would enable us to deal with this overload.
This visualization was featured on the Guardian Datablog, and presented with little explanation. When I first saw it, I was overwhelmed with a sense of complexity. If the Textexture is successful with anything, it's expressing the underlying complexity of any linear document. But to really get at the meat of what the visualization is showing, you have to jump over a few conceptual hurdles. First, we see that each word node is a different size. Intuitively, we would think that the size was a representation of how often each words was used, but not so: it is actually a measure of how many connections to different "meaning clusters" that word makes. A "meaning cluster" is designated by the color of a set of nodes. Here is a word with few connections:
Here is a word with many connections:
Second, what is the meaning behind the arrangement of the nodes? When the visualization first loads, we see the nodes dance around as if to get to the right "spot". Yet, there seems to be no underlying structure—except, perhaps, that the largest nodes appear in the center (but still, interspersed with small nodes). It could be that this is a statement in itself about the how the speech was structured. I think a short explanation in the "How to use blurb" would mitigate this. It may be that these are things we just have to get accustomed to when we encounter new ways to visualize data. Something we're not used to just need a little explanation. These things aside, this is a very neat and responsive exploration tool, and it can be used to explore the underlying connections in any written document that interests us.












