Researching, Teaching and Learning with Internet Data (part 1) : Media Cloud as an Open Research Infrastructure

Media Cloud is an open research infrastructure and archive for studying the flow of news stories on the Internet. It has been developed since 2009 by an interdisciplinary team of researchers, journalists and technologists, first at the Berkman Klein Center for Internet and Society at Harvard University, and later at MIT and Northeastern University. I learned about it during my summer internship at the Berkman Klein Center back in 2011 and became fascinated by the opportunities that collecting, processing, and accessing Internet data offered for social science and humanities research. At that time, the platform integrated a range of computational tools to collect, organize and analyze digital news content and visualize hyperlink networks, and made them available online for everyone.

Tools available in Media Cloud circa 2011-2016

Around 2011, I was mainly doing ethnographic projects in Austin and had started to experiment with the intersections of doing fieldwork online and offline. As I was advancing the fieldwork I became increasingly aware of the importance of accessing and collecting public Internet data such as the one people is openly publishing on the World Wide Web. Internet data is critical for understanding people’s everyday media practices and the range of social, cultural and economic phenomena that emerge from mediated digital interactions. This kind of data encompasses the diverse digital traces produced through online activity, with user-generated content (including posts, comments, memes, images, videos, reviews, web pages and other forms of expression) constituting a particularly valuable source for studying communication, social interaction, learning and collective behavior.

Years later, when I returned to the Berkman Klein Center as a postdoctoral fellow in 2015, the platform’s development had moved from Harvard to the MIT Media Lab’s Center for Civic Media, and the research infrastructure has become global allowing the collection of news stories from different languages. It was then when I decided to start using Media Cloud with the motivation of researching and understanding Latin American media ecosystems, and, particularly, the Colombian one. Around that time the Colombian peace process was advancing in the midst of a polarized public debate and the rise of the information disorder online. The increasing circulation of mis- and dis-information about the peace process and the 2016 plebiscite had similarities with what was happening on the US 2016 presidential election campaign, and the UK Brexit. Although the scale of the information disorder and the number of actors involved were different in the Colombian media ecosystem, some of the dynamics and tactics of problematic information production and circulation resembled what happened in the US and the UK. Highly polarized political debates created fertile conditions for misleading, emotionally charged, and divisive content to circulate rapidly, often reinforcing pre-existing identities and grievances rather than encouraging deliberation

Members of the Media Cloud research team introduced me to the inner workings of the platform and allowed me to contribute to the Colombian news sources collection. They offered me support for the kind of research I wanted to do with Colombian news data and introduced me to a range of tools for mixing qualitative, quantitative and computational methods, and for collecting other kinds of Internet data such as the digital traces created on social media platforms like Twitter, YouTube and Facebook. Thanks to this kind of exposure, I discovered the power of using social network analysis, a methodology that has become central in my toolkit.

When I returned to Colombia in 2019, several of the projects I developed leveraged the news archive of Media Cloud for researching the national and regional media ecosystems and understanding complex phenomena such as the patterns of information diffusion, political controversies and polarization. Media Cloud allowed us to build extensive corpora to systematically examine large volumes of content over extended periods, making it possible to identify patterns that may be difficult to detect through traditional, small-scale sampling. These projects functioned as a training site for developing data science skills that are critical for understanding the dynamics of networked political communication and other complex phenomena that occur in contemporary media ecosystems.

In the past five years Media Cloud has undergone a substantial technological and methodological reconstruction. After the archive surpassed one billion stories, scaling problems led the team to rebuild both the backend database and frontend interface. In 2021, its current governance structure took shape as a consortium involving the University of Massachusetts Amherst, Northeastern University, and the Media Ecosystems Analysis Group (MEAG), establishing a more distributed institutional model for managing the platform, datasets, and research infrastructure. In 2023, the project migrated to a new database infrastructure with updated methods for searching, content extraction, date estimation, and deduplication. The most significant development has been Media Cloud 2.0, which completely re-engineered data collection, storage, and retrieval, introduced new research interfaces including a searchable Story Index and a global Directory of news sources, and reprocessed the historical archive using consistent, modern metadata extraction methods. Although this changes have made the platform more robust, it has also limited the access to analytical tools such as network visualizations and the possibility to download the complete datasets that result from the queries. In its current deployment, researchers can do queries in many languages and search the news archive, but the results on the platform only display a random sample of the news stories that match the searches.

Leave a Comment

Your email address will not be published. Required fields are marked *