2017 UK edition
The Lowry, Salford · 26-27 January 2017
The first UK edition, and the start of the European series. Sixty-eight people were invited to The Lowry in Salford for a full day of presentations, a dinner, and a half day of deeper discussion. Talks ran fifteen minutes, with an equal slot for argument afterwards. The room mixed public and commercial broadcasters, media agencies, record labels, universities and analytics vendors. Two themes ran through both days: how to break content down into machine-readable attributes, and how far a company's own platform data can be trusted to describe its market.
Talks
- “Introductions and Opening Remarks: The State of Data Science and Analytics in Entertainment” – The organisers, from a public broadcaster and its commercial arm, opened the two days.
- “Content Genome 1: TV Show Genomes to Predict Country-by-Country Success” – A business-school researcher and a script-analytics firm broke 500 seasons of a broadcaster's catalogue into roughly 70 attributes, then modelled licensing revenue across 20 countries. The aim was not to pick hits but to say which themes over-index where, which is a licensing negotiation rather than a commissioning decision.
- “Content Genome 2: Hacking Together Film Data and Lessons Learned Therefrom” – A media agency stitched three years of weekly ticket sales across 11 countries to film identifiers, search volumes and trailer views. The lesson was about access rather than method: one supplier released a partially blinded sample, one title in ten, so the modelling could be checked without handing over the catalogue.
- “Content Genome 3: Metadata Extraction from TV Shows and Lessons Learned” – A broadcaster's research department ran speech-to-text and entity extraction over 70,000 archive radio programmes that had almost no metadata, then let the public correct the output. Combining subtitles, face recognition and speaker identification narrowed a search far more than any one signal alone.
- “What Makes a Personalised Weekly Playlist So Good?”
- “Using Agent-Based Modelling to Predict the Future of UK TV Consumption Behaviours” – A simulation firm showed a proof of concept: 600 synthetic agents, each needing a blend of five content "nutrients", choosing between four channels. The framing came from nutritional geometry, where an animal forages for a combination of foods that no single food provides.
- “Clustering 1: The Power of Clustering Survey and First-Party Data Sets” – A social-graph analytics firm split the followers of one television brand into 20 clusters by who they follow rather than what they post. Only 10 to 15% of users on that platform post at all, so listening tools describe a small and unrepresentative minority.
- “Clustering 2: Behaviour-Based Genres and Genre-Based Segments” – A media agency derived 13 genres from the viewing behaviour of roughly ten million people, using cleaned-up media file names as the content signal. The clusters were handed over unnamed so that the people who had to use them could name them.
- “Clustering 3: The Science of Nudging Viewers Using TV Show Profiling” – Behavioural scientists predicted whether a show would land in the top tail of public ratings with about 70% accuracy using three attributes only: genre, country of origin and the number of countries it had been released in.
- “Box Office Forecasting and the Release Date Optimisation Problem”
- “Centaur Evaluations: Merging Mind and Machine for Better Recommendations” – A broadcaster tested recommendation models by having staff judge head-to-head comparisons, with model identities hidden and the order randomised to strip out framing effects. The proposed upper bound was not a better algorithm but a panel of human editors.
- “What Affects Your Pay? What Drives Salary Increases?”
- “One Second Reach: User-Level Data and Second-by-Second Viewing Data”
- “Survivor Bias” – An analyst used the wartime bomber armour problem to argue that first-party data describes only the people who came back. Even a heavy user gives any one platform a fraction of their total time with that medium, and the rest of it is invisible from the inside.
- “Real-Time Media Analysis”
- “Why Data Science Projects Fail”
- “Forward-Looking Media Trends” – A research firm argued that subscription video was adding to pay television rather than replacing it, with nearly three quarters of a billion of combined audience growth. Messaging platforms had 6.5 billion monthly active accounts, close to twice the number of connected people, and almost no measurement.
- “The Data Behind Influencers” – A creator-search platform had confirmed more than 20,000 brand collaborations and detected roughly 200,000 more. Against the investor view that influencer marketing was finished, the like-to-dislike ratio on branded videos had climbed for a decade to about 96%.
What we learned
Politics Does Not Travel: Modelling 500 seasons of one broadcaster's catalogue against licensing revenue in 20 countries produced one of its strongest negative coefficients on political themes in the United States. That was the point of including it. A known industry fact appearing cleanly in the data was used to earn trust in the surprises that came out alongside it.
The Data Was the Hard Part: Assembling a film dataset took weekly ticket sales for three years across 11 countries, a chain of identifier lookups, scraped trailer views and two search interfaces played off each other to recover real numbers from indexed ones. What unlocked it was not statistics. It was a partially blinded sample of one title in ten, data destruction agreements, and being easy to deal with.
Armour the Planes That Did Not Come Back: First-party data feels safe because it is deterministic and proprietary. It is also a survivor sample. Even a heavy user gives any single platform only part of their relationship with music or television, and everything outside that share is invisible from the inside. The recommended fix was to anchor platform data to a probabilistic market survey before segmenting it.
Three Variables, Seventy Per Cent: Genre, country of origin and the number of countries a show had been released in predicted whether it sat in the top tail of public ratings about 70% of the time. The room pushed back hard on the third variable, on the grounds that release breadth is itself a consequence of success rather than a decision input. That argument was the session.
Editors as the Upper Bound: One broadcaster framed recommendation quality with two benchmarks: random is the floor, and the average judgement of five content editors is the ceiling worth chasing. The wider point landed harder. The recommender had several goals, most of them never written down, which makes it an optimisation problem nobody had stated, let alone solved.
Eighty-Five Per Cent Say Nothing: On one large social platform, only 10 to 15% of users post, and most of those post rarely. Social listening therefore samples a talkative minority. Clustering people by the accounts they choose to follow, rather than by what they say, covers the silent majority and produces segments that ordinary analytics cannot see.
A phrase kept recurring across the two days: the poor man's version. Nobody in the room had the data the platforms had, so every method on show was built to work without it.
