2023 LA edition
Marina del Rey · 3-4 August 2023
Twelve talks covering content valuation, streaming forecasting, genomic targeting, release window optimisation, playlist economics, and the emergence of language models in entertainment analytics. A live demonstration showed a non-programmer running a full clustering analysis of music industry data through natural language conversation.
Talks
- “The Editor vs. the Algorithm” – A field experiment at a major news outlet found the algorithm outperformed human editors, but a combination of both would have delivered 13% more clicks than either alone. Pure algorithmic personalisation also reduced consumption diversity over five months, creating a feedback loop that narrows the content a platform can justify producing.
- “Estimating Incremental Acquisition of Content Launches” – An econometric model valued individual titles within a streaming bundle by measuring their effect on subscriber churn. Tentpole releases increased total platform consumption by roughly 6%, and subscribers whose consumption dropped below their personal average were disproportionately likely to cancel.
- “Forecasting Long-Term Streaming of Emerging Artists” – A record label’s data science team trained their model on 7,000 songs with eight years of data. One case study: a song released in 2019 doing 16,000 streams per week jumped to five million per week by May 2022. TikTok virality has decoupled a song’s discovery moment from its release date.
- “Combining AI and Storytelling Expertise to Describe Content” – A human-coded dataset of clearly defined traits was used to train models that predict complex movie traits with high accuracy. Identifying trait combinations proved more valuable than examining genres alone.
- “Measuring the Value of a Title in a Subscription Bundle” – A structural demand model allocated tens of billions of dollars of subscription revenue across bundled services. The Shapley value method was abandoned because removing a single benefit entirely produced such extreme counterfactuals that results were not credible.
- “Large Language Models Will Revolutionise Entertainment Analytics” – A live demo showed a non-coder using natural language to run a full clustering analysis across dozens of countries, complete with named segments, marketing recommendations, and visualisations. Of the room, only about six attendees had corporate-endorsed language model access, while at least five admitted to using unapproved tools.
- “The Effect of Linear Television Airing on Digital Channels” – Airings of movies on linear TV produced a small but positive lift in streaming sign-ups. Yet co-licensing films to similar streaming services caused cannibalization of viewer engagement.
- “Enhancing Advertisements Through Genomic Targeting” – Content genome data was applied to film and television advertising targeting, matching ad creative attributes to audience taste profiles at scale.
- “Optimising Release Windows Using Structural Equation Modelling” – A structural model examined how consumers choose between theatrical, home video, and streaming formats, extended to include piracy, allowing simulation of alternative release scenarios.
- “How Playlists Affect Off-Platform Behaviour” – Playlisting emerging artists led to a noteworthy increase in both on-platform activity and off-platform live concert bookings. The persistent effects suggest playlisting is a potent tool for amplifying emerging careers.
- “An Insights-Driven Narrative to Support Talent Initiatives” – Data-driven approaches to talent representation decisions, quantifying the commercial and cultural impact of diversity initiatives in the entertainment industry.
- “Entertainment Economics” – An examination of economic models in entertainment, covering pricing, bundling, and consumer behaviour across digital and physical formats.
What we learned
Algorithms Beat Humans, But Not By Much: A field experiment found algorithms outperform human editors at news curation, but the effect was modest. A combination of both would have delivered 13% more clicks than either alone. When a bug caused the algorithm’s data to go stale for one week, the human editor outperformed it. Humans also won during breaking news events.
Nobody Predicted Barbie’s $162 Million Opening: Every forecast model failed to predict Barbie’s opening weekend. The concern: if teams increasingly rely on language models trained on historical data, they will systematically produce lowest-common-denominator ideas. The models work within the box and cannot get out of it.
Songs Can Sleep for Years, Then Explode: A song released in 2019 doing 16,000 streams per week jumped to 350,000 by late March 2022 and five million by May. The forecasting team had to redefine “week one” as the week streaming growth first hit 10% of its eventual maximum, not the release date. TikTok virality has decoupled discovery from release.
Shadow AI Before the Rules Arrived: A show of hands revealed roughly six attendees from major entertainment companies were using non-endorsed AI tools at work, tools they knew would get them in trouble. Only a similar number had any corporate-approved way to use a language model.
The $60,000 Certification Problem: A graduate student learned everything he needed about machine learning from a free online resource. As one panellist put it: “The information is free, and the certification costs $60,000 a year. That can’t be sustainable.” Language models may accelerate the unbundling of education from credentialing.
