Data, digital methods and mapping social complexity http://digitalmethods-seminar.org Visualizing social and semantic dynamics in the social sciences Tue, 23 Sep 2014 17:26:57 +0000 en-US hourly 1 http://wordpress.org/?v=3.8.10 Q&A: Donato Ricci w/Pedro Miguel Cruz on being a “data visualizer” http://digitalmethods-seminar.org/qa-between-donato-ricci-and-pedro-miguel-cruz/ http://digitalmethods-seminar.org/qa-between-donato-ricci-and-pedro-miguel-cruz/#comments Fri, 22 Aug 2014 16:53:02 +0000 http://digitalmethods-seminar.org/?p=726 […]]]> Donato Ricci spoke with Pedro Miguel Cruz about his conception and practice of data visualization.This interview was conducted as an extension of seminar session #6, held on 15.05.14, where Cruz spoke about “Visualizing Complexity.”

1

Donato Ricci [DR]: Would you describe yourself as a computer scientist or a designer? Do you prefer one to the other? If not, why not? Or perhaps you consider yourself a storyteller?

Pedro Miguel Cruz [PMC]: That goes into the core definition of infoviz as an intersection of several fields. I started in Physics Engineering so I had some training in dealing in solving analytical problems with analytical tools. But in the end, what I wanted was to solve concrete problems, with concrete tools and solutions. Concrete in the sense that I can point to an image and tell you “that’s the solution”. In informatics engineering, my background is more on an engineering perspective of computer science, so I cannot say that I’m a computer scientist. I often use solutions of computer science problems when building a visualization, but I rarely innovate upon those solutions (think about standard algorithms or computer graphics techniques – nevertheless I already used some solutions that could be considered new from a technical perspective in computer graphics and scientific visualization, but I don’t invest in having those solutions working as general purpose solutions outside of the context that they were created for). So I see my reasonable knowledge in computer science and graphics more as a tool. Where I innovate, maybe, is in the amalgam of techniques that I bring together to problems in different domains where they aren’t generally used.

I think it’s fair to be called a storyteller. The story comes from several choices that you make. First of all, if I’m working on a subject of my own choosing (rather than as a designer on another person’s project) then I’m making very direct editorial decisions about which aspects or issues of the subject I consider a priority to treat and treat in a way that best brings that issue to life for the reader/viewer. This involves decisions about which data dimensions to highlight and what’s the emphasis to give to these dimensions by assigning them visual representations. Furthermore the story is carried by the rhetoric of the visualization; mainly the designer can push a more normative or explicit discourse than might otherwise be present through more standard approaches. Furthermore I also make very concrete choices on an execution level in order to render such stories and discourses very clear. And in this sense design plays its role as a process of message clarification. You can say that I’m a designer and a storyteller. You can say that I’m a designer of stories (because the story is the shape that I give to a series of “events”, the data).

DR: How do you explain what you do to your family?

PMC: This is funny because if I go with “I do information visualization” they don’t have a clue. Even talking about data, design, discourses… But if I address what I do by talking about social or historical matters, and then tell them that I can “show” them so they can better understand, they pretty much get it (of course after showing the examples). They understood my projects enough that they could start asking really tough questions about them and their construction. And that, I never thought it would happen.

2

DR: Could we say that part of visualizers’ work is to observe and describe the shape of social phenomena? Do you have any particular strategy for exploring different visualization solutions that lead you to the most compelling solution? Is there always “one” solution that is the most compelling and how do you decide?

For example, the choice to disentangle a network from a graph to better grasp the data? In other word how do you describe your way of working? Is it linked somehow to what Moritz Stefaner labels as bootstrapping?

PMC: Finding a shape to social or other phenomena, giving it form, is definitely part of a visualizer’s work. The strategy that I use depends on the nature of the project. If it has very concrete objectives defined, then there are often solutions used in the same domain that work pretty well and can be used. Or often you just merge several solutions in the same visualization in order to have complementary views of the data. On the contrary, if the project is on a subject of my own choosing, I actively avoid using an already seen approach in visualization to that problem, and for that matter, I also try not to make the core of the visualization about an analytical standardized graphical strategy to depict the dataset in general. As for if it is the most compelling that is always an open debate. Even when choosing from a catalogue of standardized graphical methods, one can argue that you have several that are equally effective. When working again on your own material rather than an external assignment, well I can at least say that the solution I choose is the one most compelling to me.

I’m not familiar with the bootstrapping strategy from Moritz, so I’m not sure I completely get it, but if it refers to building a set of basic visualizations in order to help you figure out the data and then work on a next step in order to build a more elaborate, compelling one, then for sure. It’s part of the visualization building process to have already at least basic static figures that illustrate the data for you and point in directions where you could extract the most valuable narratives from data, or point in a direction where you know that the visualization model that you are thinking will accommodate that same dataset (for example, you have several solutions that won’t work with a very high number of data points, so you have to study how many you are displaying at the same time, or if the extremely elaborate solution that you came up with to depict relations among data points won’t make the visualization unrenderable–and unreadable for that matter–, from a computer processing perspective, for instance).

This brings me to another matter: details, execution and aesthetics. That is for me of the utmost importance, to work every detail of the visualization with the same care that you would if you were tailoring a suit—animations should be smooth, transitions should feel natural, it should run smoothly in realtime, if it’s 30fps, it’s 30fps, not 15, there is a layout, things should be aligned, choosing the typography is also crucial since it sets the tone of the visualization, there is a typographical grid; color theory exists and it should be used, etc. On the aesthetics side, I’m talking about clarity and elegance. Not every color is to be used in every visualization, neither every shape. I look for two things: either to show clear patterns from visual complexity as if we had a synthesized unique visual form that portrays such patterns, but where we can search for each of its tiny single constituent parts (I often feel that I’m working with textures and patterns, and trying to emphasizes differences in their density for example); or by applying Swiss attention to complexity–it’s grid-based, it’s ordered, it’s elementary in the shapes, it’s so much easier to achieve visual elegance through it.

But let me tell you a bit more about my approach: I research, research, research, because I’m always interested in making something new! I also rarely use straightforward approaches. Let’s put some salt in all those boring graphs, please. Even with network visualizations, I find myself half asleep because the graph itself is already an abstraction, and they so often have that same hairy ball aesthetics… I try to make the visualizations closer to what I think the imaginary representation people intuitively have for a given dataset. That’s my favorite approach: e.g. if you have a person traveling from A to B, I don’t draw a static line from A to B. I make a person actually move from A to B. This takes me to the core of my approach: I try to build systems that react to data. You have agents/actors in the visualization and they react to data, while having other properties and constraints that are not data related, but more related to that imaginary that I previously mentioned, and if those additional behaviors imply that they don’t always portray the data exactly the same way, well, that just makes things more interesting. I create agents and behaviors, and I set them free, and there they are feeding or their little world of data while enabling us to look at that same data just by observing their behavior. In this sense I often say that my approaches are nature-inspired.

3

DR: More and more trans-disciplinary teams are set up to observe and explore complex social phenomena. It is my impression that the more designers that are involved in these teams the more they learn more about the analysis than about data visualization. Furthermore, the more they are side by side with other disciplines the more the visualization become standardized so one could ask: “in data visualization, is innovation possible?” So then, what is innovation in data visualization?

PMC: Ahah! That is so true. When you have a large team, normalization reigns. Because there are so many insights that you want to guarantee that you can’t with only a strong graphical concrete approach–in the end you are left you with a series of standardized ones that can be used to decode all those discourses that were inputed by that large team during the process, so that everyone is happy. Furthermore trans-diciplinary teams are often not very keen in very bold approaches, or approaches that they are not familiar with.

Of course there can be innovation in dataviz! But you need a vision. And you need to choose if you are doing business reports, or visualization that is to be consumed by a very large audience in a short amount of time while engaging them. Naturally you cannot convey every vicissitude in data, but you can open a door for the awareness of their presence or directly to their exploration. There should be a vision, that may or not work, that may or not change along the way with the team’s input, but it shouldn’t lose its essence. And I’m talking about a vision because that’s the thing that can create some room to innovate, or innovation will be dragged down by the mighty forces of normalization. When innovating you take some risks, you are not sure that it will work, and you are well aware that some of the chosen solutions are not the best from an analytical perspective. But in the end: did you make the dataset more interesting than the dataset itself? Did you engage a large audience and create awareness of that dataset? For the “what is innovation in dataviz” question, you can look at it from several ways. For example, all those ways that exist to make a network layout. They are there, but they have been a long time ago, but only recently we see them more synthesized. They are more often not new, than just freshly cooked. Other times they were just buried in the past, but they were not new. But their application to a new domain problem? That might be new and that might constitute the innovation per se. I hope that in the near future we will be able to naturally address this question just like the way that we do for “what is innovation in graphic design?”. It’s about history, applications to concrete problems and philosophies. Now, we can also talk about approaches to data visualization, as tendencies that shape the field and generate what we perceive as innovation in data visualization. For example, you should have already noticed that I try to make my visualizations ludic and playful in order to engage the public. But this approach has clear gamification roots, that I haven’t yet fully implemented. Full featured gamification of data viz (and I can refer to very simple games) could be a major milestone in the field.

4

DR: A lot of these trans-disciplinary teams are dealing with the ‘second computational turn’ in the social and cultural researches.  Here two main empirical approaches could be identified: the Big Data as opposed to the Digital Methods.

The first one focuses on very large data-sets and often characterized as ‘data-driven’. It is mainly exploratory in orientation, tending to identify patterns where the nature of the patterns recognized may seem of secondary importance compared to a general demonstration of the potential analytic capacity of this type of research.

By contrast, Digital methods rather highlights the formatted nature of much digital data. It highlights there is no such thing as ‘raw data’. In practice, data and analysis cannot be distinguished in any easy or straightforward way. In other words, is to a large extent driven by research design (formulating good research questions, delineating of source sets, developing a narrative and findings).

Is there in your experience any evident way of presenting visually these two ways of dealing with the data, the analytical one and the interpretative one? In my experience the Big Data approach often leads to an anecdotical use of the visualization while the Digital Methods tends to “use” the visualization as argumentative devices. Does this make any sense to you?

PMC: Yes, I completely agree. You have some very poor uses of visualization (or sometimes no visualization at all) when dealing with some datasets and favoring only their analytical description (with some simple drawings, here and there…). We are talking about two different ways of working, seeming even that they employ different types of professionals. Of course that “big data approach” will typically lead to poor representations. When choosing a narrative you have to apply certain constraints, perhaps leave some aspects of the data out of the picture, and not everyone likes to do this a priori. It may also be much harder to do since data could be less structured than the one being used on the “digital methods approach”, and then it is harder to extract semantic contexts from it. What you can do is to tag the data yourself, structure it yourself. But even this will have you making certain assumptions that not everyone is comfortable doing so early in the process. Of course the visualizations in “digital methods” are richer, more visually elaborate, but less complex in the problem delineation. They say more to people. But let me tell you, sometimes you have certain subjects or data, that will always be contextualized as “big data”, and these subjects shouldn’t be treated using the “digital methods” approaches–the complexity of the visualization may actually alienate the interest of a public you are trying to reach.

5

DR: Both in crafting an anecdote and an argument, the design choices implied in doing visualization lead to the notion of the researcher as an author. How do you make this role visible in your visualization, if it is possible? How, as an author, do you make your design choices in connection with your foreseen public?

Santiago Ortiz, in one of his workshop, asked the students to find a way to visualize two numbers and they ended up with more than 40 ways, all equally valid. Is the notion of authorship as linked to a principle of responsibility the way to escape the urge to evaluate visualization for example in ergonomic terms as the cognitive charge or the speed of decoding information from their visual representation?

PMC: Of course it is possible. You have several author roles in the process of visualization, you choose a story (and extract it from data), you choose an audience and you even setup those additional visual cues that add to your discourse/story but are not necessarily in the data at all.  This is visible in the visualizations that I do because you have there clear graphical and functional constraints embrace and solutions that are purposely tailored for that dataset and hardly would work with others on different subjects.

Well, it seems you have more than 40 ways to represent two numbers, but are they equally valid considering your audience and story? If you take the semantic nature of those numbers you will see they are not. So here you are already making choices where what is left to know is if the choice of the story and audience are also yours.

As for relating authorship with the lack of scientific analysis of the perceptual effectiveness of a visualization, I do not think it should be seen like that at all. I can tell that those types of empirical evaluations are not my approach, but they are needed, and I certainly do mine based on the knowledge of others’ analysis. The importance here is not to obliterate all that we already know about rules for the perceptual decoding of information on the posture that we don’t want to compromise our “artist’s integrity”… First things first. Information first. You are not an artist, you are an author of a very well defined discourse that you want to communicate, a discourse that you logically built from a dataset, and the author of a design that clearly and effectively communicates that discourse.

—————————————

Donato Ricci is the lead designer on Bruno Latour’s AIME project at the Medialab, Sciences Po and Assistant Professor at the Politecnico di Milano’s Density Design Lab.

Pedro Miguel Cruz, is an independent designer and currently enrolled as a PhD student in the Doctoral Program for Information Science and Technology at the University of Coimbra in Portugal.

]]>
http://digitalmethods-seminar.org/qa-between-donato-ricci-and-pedro-miguel-cruz/feed/ 0
Data Processors // David Foster Wallace on Information http://digitalmethods-seminar.org/data-processors-david-foster-wallace-on-valuable-pertinent-information/ http://digitalmethods-seminar.org/data-processors-david-foster-wallace-on-valuable-pertinent-information/#comments Thu, 12 Jun 2014 17:36:25 +0000 http://digitalmethods-seminar.org/?p=671 […]]]> 2011_BowWow
[img: Atelier Bow-Wow, Miyashita Park, Tokyo, 2011, via Detail, das Architektur Portal]

‘The common misapprehension is that a messy desk is a sign of a hard worker.’

‘Get over the idea that your function here is to collect and process as much information as possible.’

‘The whole mess and disorder of the desk on the left is, in fact, due to excess information.’

‘A mess is information without value.’

‘The whole point of cleaning off a desk is to get rid of the information you don’t want and keep the information you do want.’

‘Who cares which candy wrapper is on top of which paper? Who cares which half-crumpled memo is trapped between two pages of a Revenue Ruling that pertained to a file three days back?’

‘Forget the idea that information is good.’

‘Only certain information is good.’

‘Certain as in some, not as in a hundred percent confirmed.’

‘Each file you examine in Rotes will constitute a plethora of information,’ the Personnel aide said, stressing the second syllable of plethora in a way that made Sylvanschine’s eyelids flutter.

‘Your job, in a sense, with each file is to separate the valuable pertinent information from the pointless information.’

‘And that requires criteria.’

‘A procedure.’

‘It’s a procedure for processing information.’

‘You are all, if you think about it, data processors.’

The next slide on the screen was either a foreign word or a very complex acronym, each letter in bold and also underlined.

‘Different groups and teams within groups are given slightly different criteria that help inform what to look for.’

The Personnel aide was thumbing through his laminated outline.

‘Actually there’s another example of the information thing.’

‘I think they’ve got it.’

The CTO had a way of turning one foot out perpendicular to its normal direction and tapping it furiously to signal impatience.

‘But it’s right here under the desk thing.’

‘You mean the deck of cards?’

‘The checkout line.’

They seemed to believe their mikes were off.

‘Christ.’

‘Who’d like to hear another example illustrating the idea of collecting information versus processing data?’

Cusk was feeling solid and confident, as he often did after a series of attacks had passed and his nervous system felt depleted and difficult to arouse. He felt that if he’d raised his hand and given an answer that turned out not to be correct it wouldn’t have been that big of a deal. ‘Whatever,’ he thought. The ‘whatever’ is what he often though when he was feeling jaunty and immune from attack. He had twice actually asked women out when in this cocky, extroverted, hydrotically secure mood, then later failed to show up or call at the appointed time. He actually considered turning around and saying something jaunty and ever so slightly flirtatious to the noisome Belgian swimsuit model – on the upswing, he now wanted people’s attention.

At age eight, Sylvanshine had data on his father’s liver enzymes and rate of cortical atrophy, but he didn’t know what these data meant.

‘There you are at the market while your items are being tallied. There’s an individual price for each item, obviously. It’s often right there on the item, on an adhesive tag, sometimes with the wholesale price also coded in the corner – we can talk about that some other time. At checkout, the cashier enters the price of each grocery, adds them up, appends relevant sales tax – not progressive, this is a current example – and arrives at a total, which you then pay. The point – which has more information, the total amount or the calculation of ten individual items, let’s say you had ten items in your cart in the example. The obvious answer is that the set of all the individual prices has much more information than the single number that’s the total. It’s just that most of the information is irrelevant. If you paid for each item individually, that would be one thing. But you don’t. The individual information of the individual price has value only in the context of the total; what the cashier is really doing is discarding information, which in the cashier runs through a procedure in order to arrive at the one piece of information that’s valuabe – the total, plus tax.’

‘Get rid of the layman’s idea that information is good. That the more information the better. The phone book has lots of information, but if you’re looking for a phone number, 99.9 percent of that information is just in the way.’

‘Information per se is really just a measure of disorder.’ Sylvanshine’s head popped up at this.

‘The point of a procedure is to process and reduce the information in your file to just the information that has value.’

‘There’s also the matter of using your time most efficiently. You’re not going to spend equal time on each file. You want to spend the most time on the files that look promising in terms of yielding the most net revenue.’

‘Net revenue is our term for the amount of additional revenue generated by an audit less the cost of the audit.’

‘Under the Initiative, examiners are evaluated according to both total net revenue produced and the ratio of total additional revenue produced over total cost of additional audits ordered. Whichever is the least favorable.’

‘The ratio is to keep some rube from simply filling out a Memo 20 on every file that hits his Tingle in hopes of jacking up his net.’ Cusk considered: An examiner who filed no Memo 20s ever would have a ratio of 0/0 which is infinity. But the net revenue total would, he reflected, also be 0.

‘The point is to develop and implement procedures that let you determine as quickly as possible whether a given file merits closer examination -’

‘- that closer examination itself involving some type or types of procedures blended with your own creativity and instinct for smelling a rat in the woodwork -’

‘- although at the beginning of your service, as you’re gaining experience and honing your skills, it will be natural to rely on certain tested procedures -’

‘- a lot of these will vary by group or team.’

‘Incongruities on the Master Files, for one thing. That’s pretty obvious. Disagreement of W-2s plus 1099s with stated income. Disagreement of state return with 1040 -’

‘But by how much? Below what floor do you simply let an incongruity go?’

‘These are the sorts of matters for your group orientation.’

David Foster Wallace, The Pale King, Penguin Books, 2012, p.342-345.

]]>
http://digitalmethods-seminar.org/data-processors-david-foster-wallace-on-valuable-pertinent-information/feed/ 0
BarCamp // 26 & 27.06.14 @ ENSCI http://digitalmethods-seminar.org/barcamp/ http://digitalmethods-seminar.org/barcamp/#comments Wed, 04 Jun 2014 14:53:04 +0000 http://digitalmethods-seminar.org/?p=613 […]]]> barcamp

Welcome to the program page for the seminar’s BarCamp. You will find all the details on the BarCamp, the cases and datasets we plan to treat, and ways to participate in the two-day event.

If you want to sign up right away, please click here; otherwise, read on.

I. BarCamp

A BarCamp is a participatory workshop geared toward either the development of web applications or the exploration of datasets using existing software tools and ad hoc code. We are focusing our BarCamp on the exploration of a set of pre-selected and pre-treated datasets. Our goals are two-fold: 1) to give participants of the seminar a hands-on experience crafting stories and analysis with digital data while working in collaborative teams of developers, designers and social science researchers; 2) to advance the visual treatment, or interface design, of datasets connected to existing research projects or civil society groups.

II. Subjects and Datasets

Our seminar began with a number of sessions that looked at how digital methods evolved through different research communities (scientometrics, history, literary studies, sociology and media studies) and how each community has adapted different tools to treat different types of data. We would like to follow a similar approach in the BarCamp by treating a range of “data types” which each have slightly different “digital” qualities. As a result, we have identified 3 cases that will help us in this exploration of “data types”:

  1. Observing community formation and online collaboration on Github
  2. Visualizing the financial impacts of climate risk in SEC 10-K filings
  3. Pesticides online: Debates, polemics and the greening of agriculture

Between the cases, we will have the opportunity to treat various kinds of datasets, including: dynamic, user generated data coming from a social networking-platform (Github); highly formatted textual data coming from the world of financial accounting (SEC 10-k filings); heterogenous textual data coming from a wide variety of media sources (collated through Factiva; press of pesticides), classic scientometric data related to scientific publications (from WoS; on impacts of pesticides), and web cartography data (discussion of pesticides).

The hope is that by proposing various datasets, BarCamp participants will either find a subject or “data type” that corresponds to research interests of their own. Thus, they will be able to draw analogies between the BarCamp cases and their own research. As for outcomes, ideally, after two days there will be results to share with the research and/or civil society groups connected to the datasets; results which they may mobilize for further research, analysis or communication within their own networks.

A more detailed description of each case and some initial research questions are included below.

III. Organization and Logistics

The BarCamp will last for two days from June 26-27. It will be hosted by ENSCI (Les Ateliers: École nationale supérieure de création industrielle) at 48 rue Saint Sabin, in the 11th arrondissement (Métro Chemin Vert / Métro Bréguet-Sabin), see here for Contact and location information. The schedule is as follows:

Jeudi 26 Juin

9h-9h30: Café
9h30-10h: Introduction
10h-10h45: Présentation des cas [Github, Risques climatiques, Pesticides]
10h45-11h: Formation des groupes
11h-13h: Travail en groupe [questions et taches]
13h-14h: Déjeuner
14h-18h30: Travail en groupe

Vendredi 27 juin

9h-9h30: Café
9h30-12h30: Travail en groupe
12:30-13h30: Déjeuner
13h30-16h30: Finalisation des travaux
16h30-17h: Préparation des présentations
17h-18h30: Présentation des travaux, discussion autour des suites éventuelles

The datasets for each case will be treated in advance by an “animator” who will do an initial cleaning of the data to help reduce grunt work during the BarCamp. At the start of the BarCamp, the animators will pitch the various cases and their datasets to the entire group of participants who will select with which group they would like to work. The role of the animator will be to help coordinate the work of the group, but the group should also collectively discuss and define the research questions that will motivate the two days of exploration as well as their own mode of functioning during the BarCamp.

Both the medialab and the platform CorText of IFRIS will assure the participation of a mix of their own members with skills in research, design or development. The organizers also plan to invite a limited number of external guests (issue experts, coders and designers) to ensure plenty of outside perspectives.

If you would like to attend we ask that you sign up so that we can plan accordingly for food. We also ask people to commit to attend the full two days, which will ensure the continuity  of doing this type of collaboration in such a concentrated period.

If you have any questions, fell free to send them to the organisers at contact@digitalmethods-seminar.org.

IV. The Cases

Github Observatory

Case 1: A Github Observatory

We invite participants to join us for an exciting study of Github – a social networking and workflow platform for open source software development. Launched in 2008, Github hosts over 11.7 million active coding projects, making it the largest code repository in the world. We use the Github data, queried through its API, to assess the dynamics of project formation and the porosity of open source communities present on the platform. The API gives us access to an unprecedented level of fine-grain actions by participants of the site, from project managers to mere observers of projects. It offers access to various stable entities, from projects and their portfolios of contributors to contributors and their portfolios of projects joined over time. We take Github as a large scale natural experiment with millions of traces that can be queried.

Some initial questions that we will be exploring through the BarCamp include:

  • How do these communities form? Are all contributors defined at once or do they join gradually? Can projects be defined according to particular profiles? What is the distribution of these profiles across the platform? What can we say about the economics of different types of projects?
  • We will use digital markers — collected on Github — to navigate the various questions that such a trace-rich platform makes possible. We may look at this data from the angle of the specific question that animates a current project tracking the community of Russian computer scientists (RCS) active in Github. The RCS project maps Russian coders across the world and documents for the first time the relations between domestic and foreign computer scientists. Can we detect specific patterns of community dynamics and trust dimension in the group of Russian contributors? Can we detect the influence of mobility on these two elements – community dynamics and trust? In addition, how can we visualize these relationships, seizing on the flows of data in a dynamic, real-time fashion?
  • Eventually, we hope to articulate the morphology of the Github communities – who does what, when and with whom – and the types of codes produced by these communities. Beyond the motto of “open source” what else is being shared on Github?

Climate risk & US SEC

Case 2: Interpreting climate risk in SEC disclosure reports

All companies listed on U.S. stock exchanges are required to disclose information that might help investors evaluate the value of a company’s stock. These disclosures—called 10-K filings—include information about how much debt a company holds, whether they have any pending lawsuits, and what kind of other business risks they may face (competitive, physical, reputational or regulatory). In 2010, the Securities and Exchange Commission (the federal agency which oversees these disclosures) issued interpretive guidance applicable to all U.S. traded companies on what sorts of disclosures are relevant to assessing the risks and opportunities they face from climate change. However, since companies exercise their own judgment as to what to include and where to place it in the report (which can span 100 or more pages), textual disclosures relevant to climate risk are often scattered between numerous different sections of the disclosure reports and can be highly variable from company to company in terms of quality and volume. As a result, while a few groups have begun benchmarking different industry sectors around the quality of their disclosures, it is very difficult to get a holistic picture of the state of disclosures, or effectively compare disclosures within and across industry sectors—something which might allow investors to begin incorporating climate risk more robustly into their valuation models.

Working with a network of U.S. institutional investors, CookESG Research has built a series of algorithms that identify sections of the filings related to climate risk and superficially analyses the content into four basic categories. The results of this analysis are available to the public via a web-interface allowing filtered-searches of the database. This data, despite its richness, has yet to be analyzed with network analysis or text mining software.

This case will involve looking for patterns across various components of this highly structured dataset, including:

  • Comparison of reporting standards between companies within specific industry sectors, such as the oil and gas sector, or the agricultural sector.
  • Analysis is widely variable; HESS, for instance, has very detailed disclosure, while companies such as Chevron have abysmal disclosure.
  • Many companies analyze only pieces of risk, while leaving others unspoken; one task would be to see whether across all the reports there was a way of collating a complete industry “risk profile”.
  • Network analysis of the similarity of language in disclosures between industry groups. In financial filing, boilerplate text is often developed by a small number of law firms who then spread the text around. Identify this boilerplate, its distribution and its authors.
  • Tracing disclosure mentions of specific climate legislation in the US and corporate stances regarding these bills.
  • Linking this analysis to political contributions of energy companies to specific candidates (http://dirtyenergymoney.com/), and reviewing voting records of these politicians would provide a way of testing the position of these companies in disclosure statements to the stances of their political candidates on these various bills.
  • Develop an ontology of regulatory disclosures. For instance, take the case of electric utilities; where does “stationary source emission limits” come into their analysis and definition of climate risk?

Pesticide impacts

Case 3: Pesticide in circulation – expertise, controversy and social media

The use of pesticides in agricultural production is an ongoing source of epistemic tension in how to balance the use of chemicals for crop protection versus the protection of human health and environmental health from the impact of these same chemicals. Many strands of research in the life science deal with pesticides: both on the production of molecules that target specific pest and crops, but also on the bioaccumulation of pesticides residues in food, in environment or even in human body. The famous controversy about DDT and about PoP have shown the performativity of controversies about pesticides in the policy and regulatory space. This is still going on with the recent European directive about pesticides, which force member states to enact national plans for the reduction in use of particular pesticides. Associated to this regulatory regime, the issue of pesticide use is still very active in the public sphere, and appears to be a lever for the promotion of alternative agriculture that are having more and more recognition (organic farming, biodynamic, intensive agro-ecology). The nature of the debate about pesticides is thus expanding, reframing public understanding and adherence to industrial food chain.

This case will make use of a number of datasets already collected in the perspective of an ongoing research projected directed by Marc Barbier at INRA SenS on “Pesticides online: debates, polemics, controversies and the greening of agriculture.” These datasets include online debates gathered through web crawling of different web-forums on agriculture and chemicals, scientific citation data (from WoS, Cab and Medline), and a corpus of traditional press articles on the topic (Factiva).

During this BarCamp we expect to:

  • Trace the extension of the « use of pesticides » as both a challenge and a controversial space. An initial approach will be to try and follow the evolution of scientific studies of pesticides (both production and impact) and produce a phylomemic approach of the “regime of proof” of this science.
  • We also expect to map debates about pesticide use in the regional press (using Factiva), considered as a good level of analysis to have a follow-up of debates in various regions of France and possibly localize these debates on specific map.
  • Finally it looks particularly interesting to establish a mapping of web sites or blogs, thanks to crawlers, in order to establish what are the main actors involved and the contents of discussion and contentions. The legacy and effects of documentary and film might be a good entry-point for that purpose.

V. Institutional supports

The BarCamp is an event organized with the following institutional supports.

MediaLab CorText IFRIS ESIEE Logo ensci retouché

 

 

 

 

VI. Registration

Online registration is now closed. If you have any questions, fell free to send them to the organisers at contact@digitalmethods-seminar.org.

 

]]>
http://digitalmethods-seminar.org/barcamp/feed/ 0
DMT7: Spatializing Data // 05.06.14 @ ENSCI http://digitalmethods-seminar.org/dmt7-spatializing-data-invited-speaker-jacques-levy-05-06-14-ensci/ http://digitalmethods-seminar.org/dmt7-spatializing-data-invited-speaker-jacques-levy-05-06-14-ensci/#comments Thu, 29 May 2014 13:10:02 +0000 http://digitalmethods-seminar.org/?p=594 […]]]> a-cartographic-turn-5

[img: Lévy & Chavanier, A Cartogramme: Switzerland: Referendum on Minarets, 2009]

In addition to the problem of how to graphically treat and visualize data (dealt with in the previous session), are a series of underlying questions about the metaphors, metonyms and metrics we deploy to translate the digital into the spatial (Levy, 2012). When we invoke digital “mapping” tools, or talk about “spatializing” our networks through tools such as Gephi, we are making loose references to practices and techniques of the field of cartography. The session will unpack the relationship between web cartography and traditional cartography by taking a long historical view of the evolution of the field and the uses (navigational, aesthetic, conquest) of its objects (Farinelli, 2009). We will also think through the epistemological commitments entailed in describing the activity of digital analysis and representation as “mapping” (November et al., 2010). What do we gain and lose by adhering to this term? And we will focus some attention on an increasing number of voices–many coming from the domain of geography–calling for critical data studies.

[14h00 -15h15] Collective discussion of the articles listed below

  • Ghitalla, F. (2013). Des boussoles et des territoires. L’Atelier de Cartographie. (link)
  • Plantin, J-C. (2012). D’une carte à l’autre : le potentiel heuristique de la comparaison entre graphe du web et carte géographique. In Analyser le web en Sciences Humaines et Sociales, Barats C; (dir). (link)
  • Wyly, E. (2014). The new quantitative revolution (link)
  • Dalton, C. and Thatcher, J. (2014) What does a critical data studies look like and shy do we care? (link)
  • Beyon the Geotag: situating “Big Data” and leveraging the Potential of the Geoweb (link)

Suggested Readings:

  • Lévy, J. (2012). A Cartographic Turn? Revue Électronique des Sciences Humaines et Sociales. (link)
  • Tsou, M. (2013). Mapping ideas from cyberspace to realspace: visualizing the spatial context of keywords from web page search results (link)
  • Lima, M. (2013) “Decoder les reseaux” chapitre 3 dans Cartographie des reseaux: l’art de representer la complexite (link)

[15h15-15h30] Break

[15h30-17h00] Jacques Lévy of the Ecole Polytechnique Fédérale de Lausanne (EPFL) will be the guest of this session and give a talk about the “spatial turn” in the social sciences, including the digital social sciences, and give us some elements to think critically with and about the relation of the digital and space.

The seminar is open to all. If you are interested in participating, however, please sign up on our website here.

Penser la/avec la carte – une approche métaphorique de l’espace : Jacques Lévy – Ecole Polytechnique Fédérale de Lausanne (1 of 2)

Penser la/avec la carte – une approche métaphorique de l’espace - Ecole Polytechnique Fédérale de Lausanne (2 of 2)

]]>
http://digitalmethods-seminar.org/dmt7-spatializing-data-invited-speaker-jacques-levy-05-06-14-ensci/feed/ 0
Digital Methods // 02.06.14 @ College de France http://digitalmethods-seminar.org/college-de-france-and-digital-methods-june-2/ http://digitalmethods-seminar.org/college-de-france-and-digital-methods-june-2/#comments Thu, 29 May 2014 08:50:13 +0000 http://digitalmethods-seminar.org/?p=585 […]]]> College de FranceLe College is holding its first Symposium related to the digital turn in the social sciences: “Big Data, entreprise et sciences sociales – Usages et partages des donnees numerique en masse.” The session is being organized by Pierre-Michel Menger, a professor at the College and sociologist at EHESS whose research focuses on innovation in labor markets and systems of work and production in the arts sector. The symposium includes 10 interventions on topics ranging from the interaction of emerging quantitative methodologies with more traditional methods of social sciences and the nature of the knowledge produced for policy the more we rely on sources of Big Data. One of our scientific advisors, Sylvain Parasie will be presenting, as will Dominique Boullier, a member of the medialab.

]]>
http://digitalmethods-seminar.org/college-de-france-and-digital-methods-june-2/feed/ 0
Newsrooms confronting the digital gap http://digitalmethods-seminar.org/struggling-for-relevance-newsrooms/ http://digitalmethods-seminar.org/struggling-for-relevance-newsrooms/#comments Tue, 20 May 2014 20:19:27 +0000 http://digitalmethods-seminar.org/?p=542 […]]]> 2014_Niemanlab

[img: n.a. Screen Shot, Nieman Journalism Lab,  2014]

Leaked NYTimes internal report takes an honest look at how the paper is still struggling to adapt its news-making (and packaging and selling) practices to survive the changing media consumption habits of its readers and the increasingly nimble and disruptive digital strategies of new competitors; non-traditional news entities that make better use of its existing content, or master the authority games of search engine algorithms better than the Times.

Interesting to compare this report with our previous session on the Web as Transformative and Bocowszki’s work on Argentinian newsrooms and their web-page monitoring practices. We’ve moved well beyond the web as a mere screen to survey and neutralize the scoops displayed on a competitor’s homepage into structuring problems of content related data (creating the long-tail), and analytical challenges of matching consumer data to effective business-side strategies.

]]>
http://digitalmethods-seminar.org/struggling-for-relevance-newsrooms/feed/ 0
DMT6: Visualizing complexity // 15.05.14 @ ENSCI http://digitalmethods-seminar.org/dmt6-visualizing-complexity-15-05-14-ensci/ http://digitalmethods-seminar.org/dmt6-visualizing-complexity-15-05-14-ensci/#comments Mon, 12 May 2014 11:05:00 +0000 http://digitalmethods-seminar.org/?p=514 […]]]> 2010_EmpiresDecline

[img: Pedro Miguel Cruz, Empires decline – revisited, 2010]

Early data visualizations in science ordered information in tree-like representations to address issues of classification and genealogy. The Encyclopédie’s Systême figuré des connaissances humaines and Darwin’s Tree of life are classical examples of this first period of data visualization. The recent shift towards issues of organized complexity in scientific inquiry (Weaver, 1948) has changed the practice of visualization, marking a transition from trees to networks. Despite a rich stream of research, network visualization still lacks a basic grammar of standardized graphic presentation as that advocated by Willard Brinton (Brinton, 1939) and Jacques Bertin (Bertin, 1999). The session will first address visualizing complexity from a analytical perspective, stressing the importance of information design and visual standards for improved perception and understanding of complex phenomena. It will then critically discuss what some have labeled “an established language of graphic abstraction” and present an expressive approach to data visualization.

[14h30-15h45] The seminar will start with a collective discussion of the articles listed below. There will be a brief presentation and comments on the texts provided by Débora de Carvalho Pereira and Axel Lagnau to help launch the discussion.

  • Cruz, P. & Machado, P., (2011), Generative Storytelling for Information Visualization. Computer Graphics and Applications, IEEE, 31(2), 80–85. (link)
  • Friendly, M., (2008), A Brief History of Data Visualization. In Springer Handbooks Comp.Statistics. Handbook of Data Visualization. Springer Berlin Heidelberg, 15–56. (link)
  • Healy, K. & Moody, J. (forthcoming), Data Visualization in Sociology. Anual Review of Sociology. (link)
  • Offenhuber, D., (2010), Visual Anecdote. Leonardo, 43(4), 367–374. doi:10.1162/LEON_a_00010 (link)

[15h45-16h00] Pause

[16:00-17:15] Pedro Miguel Cruz of the Universidade de Coimbra will be the guest of this session and give a talk on Storytelling, metaphors and visualization. His presentation will be discussed via Skype by Donato Ricci of Science Po’s medialab and the Politecnico di Milano. Cruz’s talk will offer a counterpoint to the articles discussed in the first part of the seminar, focusing on expressiveness rather than standards in data visualization. According to Cruz, the current rise of information visualization has breathed new life, unfortunately, into an established language of graphic abstraction (Brinton, 1939, Bertin, 1967). With the increase of publicly available information related to the social sciences (urban development, history, political science and sociology) data is often treated with abstract visual representations suitable for analytical analysis (Friendly, 2008). This session will address alternative idioms of visualization through an exploration of a number of case studies. These alternative languages of the visual are intended to produce visual applications that better acquaint broad audiences to publicly relevant datasets, create compelling data narratives and raise awareness for the subjects under study. This session will address topics such as storytelling in visualization (Cruz, 2011), the role of graphical metaphors to enrich communication and the use of imprecisions in the representation of information in order to emphasize certain aspects in the data. This presentation takes the stand that visualization should be a more expressive activity and that developers of data visualizations need to wield the plume of the storyteller as much as the protractor of the engineer.

[17:15-18:00] Open discussion

The seminar is open to all. If you are interested in participating, however, please sign up on our website here.

Storytelling, metaphors and visualization : Pedro Miguel Cruz – University of Coimbra (1 of 2)

Storytelling, metaphors and visualization : Pedro Miguel Cruz – University of Coimbra (2 of 2)

]]>
http://digitalmethods-seminar.org/dmt6-visualizing-complexity-15-05-14-ensci/feed/ 0
DMT5: Transformative interactions: web effects on social dynamics // 17.04.14 @ ENSCI http://digitalmethods-seminar.org/17-04-14-transformative-interactions-web-effects-on-social-dynamics-1400-1745-ensci/ http://digitalmethods-seminar.org/17-04-14-transformative-interactions-web-effects-on-social-dynamics-1400-1745-ensci/#comments Mon, 14 Apr 2014 12:29:36 +0000 http://digitalmethods-seminar.org/?p=468 […]]]> 2010_Blogsstreams
[img: David Chavalarias, Attention streams in the blogosphere, 2010]

This seminar will address a critical question in the application of digital methods for social science research. The web is not merely a new resource that, through the treatment of large collections of data, lets us falsify or verify long- held assumptions about the relationships between institutional culture, individual behavior and other key concepts in the social sciences. The web itself is changing the way institutions function (such as how news is produced [Bozkowski, 2009] or science gets published [Evans, 2008]), as well as how individuals interact (social networking sites offer a new forms of the presentation of self [Goffman, 1959 ; Menaker, 2013], and commentary on blogs and news sites have spaw- ned new norms in communication). What we propose to address in this seminar is not a methodological question, but an epistemological question. How does the internet itself shape social phenomenon and require new theorizing about our objects of study? We will look at examples in the production of science and the news, and the treatment of data from Facebook and blog communities.

[14:00-15:15] The seminar will start with a collective discussion of the articles listed below. There will be a brief presentation and comments on the texts provided by Anders Munk & Andreas Birkbak to help launch the discussion.

  • Boczkowski, P. J. (2009). Technology, Monitoring, and Imitation in Contemporary News Work. Communication, Culture & Critique, 2(1), 39–59. doi:10.1111/j.1753-9137.2008.01028.x (download)
  • Evans, J. A. (2008). Electronic Publication and the Narrowing of Science and Scholarship. Science, 321(5887), 395–399. doi:10.1126/science.1150473 (download)
  • Gillespie, T. (2014). The Relevance of Algorithms, in Gillespie, T., Boczkowski, P. & Foot, K., (eds.). Media technologies, Cambridge, Mass.: MIT Press. (link)
  • Menaker, D. (2013). Taking Our Selfies Seriously. The New York Times. (link)

[15:15-15:30] Pause

[15:30-16:15] David Chavalarias, Director of the Complex Systems Institute of Paris Ile-de-France, will be the first guest for this session. He will give a presentation on the Unlikely meeting between von Foerster and Snowden: when the second cybernetics gives insights on the Big Data revolution. Almost fourty years ago, the father of the second cybernetics, Heinz von Foerster, conjectured a strong relation between collective social dynamics and the nature of inter-personnal interactions. This conjecture was called “the von Foerster conjecture” by Jean-Pierre Dupuy, who turned it into a theorem with Moshe Koppel and Henri Atlan in 1987. However, the theory generated little interest in the scientific literature, probably because of the lack of adequate social data required to experiment its predictions. The data deluge stemming from the web and social networks has changed this situation. In this presentation, David Chavalarias will analyze different scientific studies which corroborate von Foerster conjecture. He will then address the Snowden revelations and their interpretation in the light of this powerful intuition.

[16:45-17:30] Vincent Lepinay, Associate Professor at Science Po’s médialab will be the second guest for this session and present an ongoing research project on mapping russian geekography. His presentation is entitled Did Russians annex GITHUB and will focus on the GITHUB platform to analyze the habits of a population that is both notorious but under studied, the Russian computer scientists. Taking advantage of structured data made available by GITHUB, the project aims to look at participation and collaboration patterns of Russians in GITHUB projects.

The seminar is open to all. If you are interested in participating, however, please sign up on our website here.

Unlikely meeting between von Foerster and Snowden: when the second cybernetics gives insights on the Big Data revolution : David Chavalarias, Director of the Complex Systems Institute of Paris Ile-de-France

]]>
http://digitalmethods-seminar.org/17-04-14-transformative-interactions-web-effects-on-social-dynamics-1400-1745-ensci/feed/ 1
Demystifying Networks // Practical introduction http://digitalmethods-seminar.org/demystifying-networks-practical-introduction/ http://digitalmethods-seminar.org/demystifying-networks-practical-introduction/#comments Fri, 11 Apr 2014 16:09:57 +0000 http://digitalmethods-seminar.org/?p=451 […]]]> 2014_Scottbot_Irregular

[img: Scott Weingart, Networks Demystified, 2014]

Network graphs and network spatialization software lurk constantly in the background of our seminar and are a core tool in the visualization of digitally derived data (think Rogers, Leydersdorff, Cointet, Lermercier and Marres). But it is not a topic that we’ve addressed head on yet.

For those who have not yet had the pleasure of tinkering around with Gephi, CoreText Manager or other software, or for those looking for a little primer on networks, here is a very helpful step-by-step introduction by Scott Weingert, from the Information Science Department at the University of Indiana, to how to think about data through networks, and how to think about when a network might or might not be useful to use with a given dataset.

More info here.

]]>
http://digitalmethods-seminar.org/demystifying-networks-practical-introduction/feed/ 0
DMT4: Natively digital data mapping // 10.04.14 @ ENSCI http://digitalmethods-seminar.org/27-03-14-dmt4-natively-digital-data-mapping-1430-1800-ensci/ http://digitalmethods-seminar.org/27-03-14-dmt4-natively-digital-data-mapping-1430-1800-ensci/#comments Fri, 04 Apr 2014 12:47:50 +0000 http://digitalmethods-seminar.org/?p=437 […]]]> Noortje_Marres_Program

[img: Noortje Marres, Mapping WCIT with Twitter: Issue and Hashtag Profiles, 2012]

Thirty years ago, the democratization of IT radically changed the way we access, generate and manage information. The Internet has amplified and accelerated this phenomenon, producing ever increasing amounts of “natively digital data” (Rogers, 2013). This has fostered numerous studies of online culture, where researchers have turned to user-populated platforms such as Twitter to detect the associative practices of novel communities, or to sites such as Wikipedia where recent studies compare the controversality of topics on different language sections of the online encyclopedia (Yasseri, 2012). Beyond these specific studies of web-based-media use, there are broader questions about what exactly are we studying when we analyze hyperlinks, online forums, websites, etc.? Furthermore, what are we doing when we access information through ranking systems provided by search engine algorithms (e.g. PageRank) that constantly evolve to take into account a user’s prior searches? The session aims to develop a reflexive understanding of using natively digital data as a resource for research.

[14h30-16h00] The seminar will start with a collective discussion of the articles (and video presentation) listed below. There will be a brief presentation and comments on the texts provided by Alexandre Hocquet and Tao Hong to help launch the discussion.

  • Marres, N., & Weltevrede, E. (2013). Scraping the Social? Journal of Cultural Economy, 6(3), 313–335. doi:10.1080/17530350.2013.772070 (link)
  • Yasseri, T., Sumi, R., Rung, A., Kornai, A., & Kertész, J. (2012). Dynamics of Conflicts in Wikipedia. PloS ONE, 7(6), e38869. doi:10.1371/journal.pone.0038869 (link)
  • Kelty, C. (2014). The fog of freedom. In Gillespie, T., Foot, K., and Boczkowski, P., editors, Media Technologies: Essays on Communication, Materiality, and Society. MIT Press, Cambridge, MA. Video link of presentation of text:  https://www.youtube.com/watch?v=ihY0exiwkn8

[16h00-16h15] Pause

[16h15-18h00] Noortje Marres, Senior Lecturer at Goldsmiths University of London will be the guest for this session and present The Ambiguity of Social Media Research: Re-mediating Science and Technology Studies. Her talk will address a distinctive problem in social media research, namely the inherent ambiguity of its object. Of much social media research the question can be asked of its practitioners: are they studying society or technology? Not just the objects, but equally the methods of social media analysis are marked by ambivalence. They are of uncertain provenance, invoking methodological traditions in qualitative and quantitative research at once, and leaving it unclear whether they derive from media culture or from social research. Drawing on work in science and technology studies, Marres will argue that this ambiguity of social media research should not be regarded as a problem-to-be-solved. Instead, the confusion creates opportunities for understanding and can be deployed to generate insight in/as social media research. One could even say that social media research fails when the ambiguity of its object and methods is solved too quickly.

The seminar is open to all. If you are interested in participating, however, please sign up on our website here.

Ambiguity os social media research, re-mediating science and technology studies : Noortje Marres – Goldsmiths University of London (1 of 2)

Ambiguity os social media research, re-mediating science and technology studies : Noortje Marres – Goldsmiths University of London (2 of 2)

]]>
http://digitalmethods-seminar.org/27-03-14-dmt4-natively-digital-data-mapping-1430-1800-ensci/feed/ 2