Wednesday, 28 February 2024

MWC 2024: 5G still awaits killer app as mobile industry looks to AI

IBC

article here

AI was the buzz at Mobile World Congress 2024, but making a return on their 5G investment was at the forefront of operators’ minds.

The mobile industry does not want to repeat the mistakes of 5G where streamers like Netflix extracted all the value from data carried over their networks.

“5G is the fastest growing mobile standard in history but there are big challenges ahead,” said Mats Granryd, Director General, GSMA, setting the agenda for the Mobile World Congress (MWC) in Barcelona. “Mobile revenue growth has gone down but capex has gone up. We have to keep investing in new infrastructure to keep the world connected. “By 2030 AI could contribute $50 trillion to the global economy. As an industry, we need to think ethically about developing and using the tech for the benefit of everyone.”

The convergence of AI with 5G (and future 6G) networks attracted the major cloud and computing providers to the event in Barcelona. McKinsey forecasts a potential $300bn market to be unlocked if operators expose their networks to APIs from the developer community and cloud providers.

The relationship is symbiotic: AI applications won’t scale without mobile networks and mobile networks need to move to a more efficient data processing where AI runs on in the cloud. In reality, that means huge data farms with thousands of servers.

Microsoft claims it is spending more money than anyone else on building AI data centre infrastructure and arrived at MWC touting recent $5bn investments in Spain and Germany.

“AI is a new sector of the economy,” said Brad Smith, Microsoft Vice Chair and President in a keynote. “AI is the most important invention for the life of the mind since the invention of the printing press.”

Like Guttenberg’s invention, “AI is a tool that helps people to think, reason, share and learn,” Smith said.

The comparison did not end there. He also noted how printing technology was distributed across the world, “exploding” the number of books published from essentially zero to 20 billion by the end of the twentieth century.

It also led to the Renaissance and fermented the spread of democracy, he said, linking that to the potential of AI. “It should inspire us when people use technology to do good for others,” he said.

AI is already everywhere

Michael Dell, the Chairman and CEO of Dell Computers, made a similar pitch for telcos to tie networks with AI running on its data centres.

“The opportunity is bigger and moving faster than anything we’ve seen before,” he said. “There will be a significant proliferation of open and closed source large language models (LLMs). There will be multimodal models, local models and globally distributed ones. Models on your phone, nano models, models on your PC. AI will be everywhere, just like the internet today.”

Deutsche Telekom has already deployed AI in 400 use cases across the company and has built an AI ‘competence centre’ of experts developing AI-based products. It is also a signatory to a pact with other operators including SK Telecom and Singtel to build an LLM for telco-specific services.

“We do not want to be dependent on hyperscalers [like Microsoft Azure and AWS],” said Deutsche Telekom CEO Tim Tim Höttges. “We want to build our own system without [help from] the outside. I think we understand our world better so we have to train the LLM ourselves.”

Calling AI a “game-changer” for this industry, Mike Fries, CEO, Liberty Global, said he was leading his company’s charge from the top.

“We are not doing this by the seat of our pants but in a thoughtful way. Getting AI right is 10% of the model, with 20% IT and 70% about people. The organisational changes required are a massive lift. We want to bring everybody along with us.”

Australian telco Telstra aims to have all of its internal processes including customer service enabled by AI within 18 months.

“Tens of millions of data points cross our networks every five minutes,” said Vicki Brady, CEO, Telstra. “Humans can’t deal with that. We believe it’s got to be a whole business strategy, not a tech strategy. Rather than testing things, we are embracing AI now. For instance, our customer service teams are using GenAI now to get information super quick to customers.”

MWC gave a keynote to Demis Hassabis, the British computer scientist and video game designer who co-founded DeepMind in 2010 and sold it to Google in 2014 for $500m after which its AI system beat the human world champion at the complex game Go.

He said companies like Google are working to develop artificial general intelligence (AGI) - the next level of AI which will be able to perform almost any cognitive task that humans can.

“The human brain is the only reference point or proof of existence we have in the universe for AGI. The goal is to mimic all the cognitive capacity that humans have.”

He said current generation AI had gaps including the ability to plan and use memory – essentially to think. “It will be a gradual process rather than a sudden step change. Will we know when we see it? It may be obvious, on the other hand, we may have to test for it, on thousands of tasks and see if it passes the AGI threshold on all those.”

Immersive future not here yet

Two years ago, it wasn’t AI but the metaverse that everyone was talking about. In 2024 the metaverse might have dropped off the face of the planet. While there is use of XR headsets in the industry, mobile consumer VR experiences are non-existent.

There actually isn’t much value for the mobile industry in consumer VR because it is all conducted over wi-fi in the home.

“There are no 5G-enabled headsets, period. Nor is there likely to be since VR happens indoors. Out of home is where it gets interesting for mobile and where we need to provide connectivity,” said Leslie Shannon, Head of Trend & Innovation Scouting, Nokia.

When it does, new standards are needed to distribute immersive video and haptic feedback. Along with audio and visual information, experiences in the future of consumer XR – which Apple badge spatial computing and is the metaverse by another name – will incorporate tactile senses.

“The immersive future is nearer than you think,” said Valerie Allie, Video Solutions Senior Director, InterDigital, citing the Las Vegas Sphere and Vision Pro. “How can we blur the line between the physical and virtual and communicate physical sensations with virtual ones?”

InterDigital is helping develop a new haptic media format within MPEG which will outline how to create, stream and render haptic content.

“Thanks to this, creatives will be able to create content with haptic media synced with audio and video,” she said.

The company was demonstrating a connected game allowing visitors to explore the tactile sensations of live game play.

“With this new standard, we can create immersive haptic content which we can stream and distribute to the latest immersive headsets for large-scale deployment.”

App developers for XR headgear like Vision Pro will be able to create personalised experiences delivered in real-time over 5G based on the user’s actual context. This had technologist Cathy Hackl worried about the data privacy implications.

“Virtual air rights will become a new battleground,” she said. “When you wear devices like this another corporation will be able to see everything around you. Maybe sell your data to deliver hyper-personalised advertising. Your senses become real estate. How do you control what you want them to see? Who owns the space in your house? I should be able to see what I want in my house but the amount of data these devices will have about you means privacy will become a really big issue.”

Mind the EU gap

A GSMA Intelligence report noted that 5G has made great strides with the technology now available in most countries but that revenue has not grown at the same pace.

“Economic weakness has played its part but a bigger issue is that price premiums are eventually competed away in the absence of a ‘killer app’ that people will pay more for,” the analyst report noted. It is business-to-business applications that remain a focus for the industry even with the launch of 5G Standalone (end-to-end 5G networks that don’t rely on previous generations of mobile tech) “underpinning growth for the next three years.”

Europe’s operators also point to being frustrated by EU regulation that has inhibited their growth. The EU even warned in January that sluggish 5G deployment risks delaying other technologies dependent on fast internet such as AI.

Although 5G has reached 80% of the EU population, it is far below the 94% in Japan and 98% in South Korea and the US. Around 40 million people in EU countries will still have no access to a fixed gigabit connection by 2030 according to the EU’s own figures, and will therefore fail to meet the block’s target of providing 5G to all households.

“European policymakers need to change direction now,” Margherita Della Valle, Group CEO, Vodafone told the conference. “They need to reboot our legacy telecoms regulation and create a real single market supporting 5G standalone at base.”

In its State of Digital Communications report, telecom lobby group ETNO warned that significant additional investment in roll-out is still needed before EU targets to reach full 5G and full gigabit coverage by the end of this decade are achieved.

“Europe needs investment and investors need to see changes – such as a new competition policy to reach the economies of scale that 5G requires,” Della Valle said. “We need one set of rules not 27 sets of national rules. Imagine if the billions we waste today through fragmentation could be invested in higher quality, less congestion, more innovation. European economies need our investment.”

Responding, the EU Commissioner for Internal Markets Thierry Breton told the conference that future spectrum auctions for 6G should not go to the highest bidder but to the operator committed to investing in the quickest network rollouts.

He said that in Europe there was a €200bn gap in funding needed to bring EU states on par with other major economies. Latency of 200 milliseconds he said “is much too long to ensure reliable monitoring, vital telemetry or to prevent accidents between connected vehicles.”

 


Monday, 26 February 2024

MWC24: AI Has Telcos Thinking About the Future of Networks

Streaming Media

AI is set to turbocharge the mobile phone industry making sweeping changes to internal processes and customer applications from telemedicine to personalized live sports broadcasts, according to executives at the Mobile World Congress, Barcelona.

article here

Artificial Intelligence was barely on the radar of the mobile industry a couple years ago. Now, every telco and seemingly every exhibitor at the world’s biggest event for mobile network operators, has AI on the brain.

“AI is set to make sweeping changes to the way telcos conduct their business,” said Mats Granryd, Director General of mobile operators’ body GSMA. “By 2030 AI could contribute $50 trillion to the global economy. As an industry we need to ethically about developing and using the tech for benefit of everyone.”

“AI is the biggest revolution in human history,” said José María Álvarez-Pallete López, chairman and CEO, Telefonica. “For the first time, humans have a technology that can think.”

To underline the point, MWC gave a keynote to Demis Hassabis, the British computer scientist and video game designer who co-founded DeepMind in 2010 and sold it to Google in 2014 for $500 million.

“The advent of AI is an incredible opportunity for mobile and telcos,” Hassabis said. “It will super-charge digital assistants into smart assistants that are actually useful in daily life.”

He questioned, though, whether the mobile phone will be as important a device in a few years as it is today. “AI will make smartphones much more intelligent, but in 5-plus years, is the phone the perfect form factor? Could it be glasses so that AI can see more of the context you are in?”

Álvarez-Pallete López predicted telcos would swiftly move from human-operated networks to one of computer operation supervised by humans. “When you deploy fibre and 5G and switch off the old technology you have something different. The networks become software-based. Then you apply AI and the networks become proactive. It is much more than a telephone network.  This is now a supercomputer we have created,” he said. “The journey to fully autonomous networks is unstoppable.”

“There’s no way we can ignore AI,” agreed Deutsche Telekom CEO Tim Hoettges. “We will see thousands of new use cases in the next few years.” Hoettges said that telcos can maximise labour productivity by up to 37% by embedding AI in their business. AI can help deliver energy efficiency savings, predictive maintenance, and individualised customer service, he said. 

Deutsche Telekom is already deploying AI in 400 use cases across the company and has built an AI “competence centre” of experts developing AI products, he said. It is also a signatory to a pact with 12 other operators including SK Telecom and Singtel to build an LLM for telco specific services.

“We do not want to be dependent on hyperscalers,” Hoettges said. “We want to build our own system without [help from] the outside. I think we understand our world better, so we have to train the LLM ourselves.”

Today, DT makes €10 billion in revenues from European B2B customers. “They desperately want an AI solution for their business,” Hoettges said. “We can be a facilitator.”

Hoettges also emphasised that at DT, AI was a “command strategy from the CEO on down” but that AI skills needed to be upgraded at every level in the organization.

Calling AI a “gamechanger” for this industry, Mike Fries, CEO, Liberty Global said AI was front and centre of his agenda. “Tim mentioned 400 use cases. We are thinking about four,” he said. “It’s impossible to go from zero to sixty and be effective. We aim to be focused,” he explained. Second, while we love to be first, it is more important to be fast. You don’t have to be first in every application or idea. We are feeding this from the top of the company. I am leading it. We are not doing this by the seat of our pants, but in a thoughtful way. We are worried about deployment but also ethics, regulation, insights, and clients.”

One of Liberty Global’s AI focuses is on customer service. Another is the huge area of network optimisation, planning, and energy saving. The biggest is upskilling the company’s staff with tools and mindset.

“Getting AI right is 10% the model, 20% IT, and 70% about people,” Fries said. “The organisational changes required is a massive lift. We want to being everybody along with us.”

Telstra’s ambition is to be the leading “AI-fuelled organisation in Australia,” said Vicki Brady, CEO, Telstra. “Practically, how do we make it happen? Tens of millions of data points cross out networks every five minutes. Humans can’t deal with that. At Telstra we believe it’s got to be a whole business strategy not a tech strategy. Rather than testing things we are embracing AI now. For instance, our customer service teams are using GenAI now get information super quick to customers.”

She said half of Telstra’s processes today are enabled by AI with the goal of 100 percent by July 2025.

While some execs were pondering the vast sums their companies could make by monetizing new AI-driven services, few voiced concerns about its environmental cost. Álvarez-Pallete López said a single AI query “costs the equivalent in power consumption to create one bottle of water or 2 hours of an electric light bulb.” With AI and 5G, he said, “data volumes will rise four-fold between now and 2030 with huge power implications. Massive data from sensors and massive processing capacity presents new challenges including the responsibility of telcos to enhance sustainability.”

 


Friday, 23 February 2024

Spain & Portugal: Sunny outlook for uncertain times

AV Magazine

article here

As reliant on tourism, retail and live events as much of Iberia is, the region is still unwinding from the effects of the pandemic but growth should be strong moving forward due to both countries haven taken advantage of European funds.

Iberia has been more buoyant than other European markets over the past 24 months, reports Cecilia Wills, country manager, Matrox Video: “The demand for pro AV equipment and services has increased in various sectors from corporate and events to education.”

“Spain and Portugal are still existing in a post-Covid period as many smaller projects are being fulfilled that were put on hold and held in backlog,” says Jeroen Helms, sales director, EMEA, Peerless-AV. “Stock levels and logistical issues are now solved, making this fulfilment possible. New, larger projects are not as widely available in these territories, and both countries are extremely cost conscious with price driven brands dominating the pro AV market.”

Julián Oltra, MD of Clear-Com partner, Audio Video Zentralmedia also notes both the large number of projects accumulated due to the hiatus as well as the business this means for almost all AV companies.

Both countries have taken advantage of the Recovery, Transformation and Resilience Plan financed with European funds. “This will represent an economic boost aimed at the digital transformation of public and private sectors, as well as the creation of digital, safe and sustainable infrastructures,” says Carmen Jerez, regional sales manager of Southern Europe, Datapath. The investment is worth 20 billion Euros over three years “a volume of resources destined for digitalisation with a scope and magnitude that will allow a truly transformative impact and reflected already in the growth of opportunities for AV.”

The global macroeconomic situation weighs on the business climate too albeit that the outlook for 2024 is for positive growth and lower inflation.

Pedro Ballesteros, Kramer’s regional sales director believes instability due to the conflicts in Ukraine and Palestine has impacted business, as well as inflation.

“These combined to result in negative growth of the pro AV market and as a consequence investments in pro AV solutions fell during the year.”

Jerez counters that despite the conflicts creating uncertainty, “it is important to note that they are two growing markets in the AV industry and especially in the control room environment,” Jerez says.

Key to navigating these uncertainties, adds Helms “is an open mind and a willingness to adapt and learn.”

“The climate is very positive,” asserts Antonio Ortega, Creston’s country manager who is especially eager to see what will happen to the industry now that the supply chain issues are under control.

Portuguese AV market
The Portuguese AV market, in particular, has been experiencing “significant and sustainable growth” reports Miguel Dominguez who looks after Genelec there.

There are no major differences between the two countries from an AV standpoint. Spain is the far larger and more developed market, with a more established and developed channel of AV distribution partners, notes Helms. “The Portuguese tend to be more protective of their smaller, more independent AV market, with a we-can-do-ourselves attitude,” he says.

Both countries have been earlier adopters of many new trends, says Ortega. “The customers in this region are tech savvy. They know what they want and push us to use the full potential of our products. Traditional products are less in demand. AV-over-IP for instance, has taken over as the most used solution, leaving the traditional solutions behind.”

Manuel Rivera, sales and marketing manager of Yamaha Music’s Pro Audio Division says: “Both nations uniquely combine historic and modern elements in their audio visual sectors, distinguishing them on the European scene as centres where cultural heritage meets technological advance.” That’s evident from last year’s tally of visitors from the region to ISE in Barcelona.

There is another notable attraction common to both countries. “Many multinationals are creating availability zones for data centres on Iberian soil, making it one of the emerging markets and also accommodating the AV industry mainly through control rooms,” says Jerez.

There are currently 60 data centres under construction in Spain alone. Amazon, Google and Microsoft are investing millions of Euros to build capacity there. At Sinces, south Portugal, Datapath is involved in building a control room for one of the largest and greenest data centres in Europe.

“The availability of cheap local green energy combined with geographical proximity to three other continents with fast connections using new high-speed subsea cables which crosses the Atlantic to Brazil, makes Sines an ideal location,” explains Jerez. “It will propel Portugal on to the international data traffic and data centre stage.”

Compared to other European markets, Spain and Portugal can be more price-sensitive for, say, corporate and education projects, says Wills. “But price is less important in mission-critical deployments, such as rail and air infrastructure projects, where brand reputation is more important.”

There’s also been an increase in public tenders in transportation due to investments in national infrastructure. Wills believes this is driving demand for video walls, AV-over-IP solutions, and IP KVM. Some of these programmes are longer-term and won’t deploy for another few years, but others are looking to close in 2024.

“Spain and Portugal are also looking towards smart cities in order to improve traffic management, citizen security, and energy efficiency,” she says. “Pro AV has a big role to play here and it will be an exciting growth area in the region.”

Tourism back to the top
Tourism is the main vertical and source of income, especially in Spain and Southern Portugal. “With the post-Covid return of tourism, the importance of customer experience has increased dramatically,” says Donald De Witte, regional director, Lightware Visual Engineering. “The list of top music festivals is increasing every year which has an impact on the live events industry. Rental and staging is returning to pre-Covid levels and all concerts are increasing their use of video and imagery. During the summer, almost all towns hold celebrations. To attract foreign visitors they are looking for more impactful events.”

Similarly, the retail and hospitality markets are thriving with more interactive experiences. The number of flagship stores in the main cities has doubled and digital signage usage has increased, says De Witte. “The presence of interactive kiosks is increasingly common.”

The amount of AV technology going into high street stores is incredible, exclaims Helms. “And where a large display is needed, this means an opportunity to attach a mount! In retail, Spain is a test ground, and we will likely see it moving into other European countries very soon.”

Crestron notes an uptake in major luxury hospitality, retail and leisure projects, including meeting rooms and guest experience in VIP rooms. “It’s clear that the focus on the user experience is growing,” says Ortega. “Ever since lockdown, there’s a heightened expectation from guests about their in-room entertainment, and end users are responding to those requests.”

Genelec has scored considerable success since teaming with sales partner Garrett. Its first major project was JNcQUOI, combining a high-class restaurant with an upmarket fashion store and delicatessen bar, distributed over three floors in the most exclusive area of Lisbon. Others followed including restaurants as well as retail stores, universities, museums and even training centres, such as the L’Oreal premises located in Lisbon.

Although longer term, it is worth considering the business opportunities of the FIFA World Cup 2030 because it involves Spain, Portugal and Morocco. “The demands on all sports facilities and broadcast equipment will be very high,” says Oltra. “Every time we see more surprising stadiums as an example of architecture, but also think about diversifying their applications with first-class connections and communications.”

Prosegur’s iSOC in Madrid
For monitoring and threat management, operators at Prosegur Security’s new intelligent security operations centre (iSOC) in Madrid rely on a video wall powered by Matrox D-Series graphics cards.

The video wall pulls information from more than 300,000 sources that the iSOC monitors — including CCTV, biometric ID applications; panic buttons and alarms and the mobile communication systems of every one of the security guards.

The iSOC is capable of remotely controlling 120,000+ cameras in more than 25,000 facilities in any sector from retail shopping centres to banking networks. AI specialised for security — developed or adapted by Prosegur in cameras and sensors — generates alerts when it anticipates a risk or a suspicious pattern.

The technology manages 1.2 million alarms a month and can discriminate false alarms, meaning that humans must manage only 1% of them. More than 1,200 specially trained operators maintain realtime contact with 26,500 security guards and technicians across Spain from this control centre.

Education sector growth
In Kramer’s case, the educational sector, specifically the university and postgraduate environment, has seen very important growth. Here, collaboration solutions for students such as VIA, or tools for teachers with QuickLaunch for education, have provided “very simple interfaces, facilitating their day-to-day tasks with a smooth transition to these new solutions,” says Ballesteros.

“In the corporate market post-pandemic, office occupancy dropped a lot and is one of the reasons for the decrease in the level of investment in pro AV technologies in this vertical, but they still need collaboration, hybrid communication and to increase the employee efficiency from the office or from home.”

More than one respondent to AV identifies the need to make personal connections as a key to success. “Being present regularly in the market and having face-to-face meetings with partners and customers is crucial,” says Wills.

Jerez identifies the “differential value of support and personalised attention, highly valued in both cultures,” while Oltra talks of a market that forces him to continually learn. “It is very stimulating for all of us who are lucky enough to work in this sector.”

“In Spain and Portugal, face-to-face interaction is even more important than normal,” agrees Helms. “Cementing business discussions over lunch or a glass of wine is expected, as a sign of obligation and loyalty.”
Salud/saúde to that.

 


AI Solutions for Climbing (Captioning) the Content Mountain

NAB

Transcription is one process that stands to be uniquely impacted by recent developments in AI. Thanks to ever-evolving language and learning models, transcribing audio to text has never been faster or easier. But there are also limitations to new AI-powered transcription solutions.

article here

The global translation service market will exceed $47 billion by 2031, largely driven by media and entertainment. Yet current costs to caption titles for distribution on streaming services ranges between $60-$100 per program hour, and typically takes between 1-3 days to complete “because of excessive manual intervention” claims Cineverse CTO Tony Huidor.

“Captions, and localization more broadly, are generally major pain points for content owners seeking to monetize their assets across the many streaming services,” Huidor added.

That’s because content companies need to generate far more revenue by broadening their audiences at significantly reduced costs.

“Companies have been priced out of bringing their entire content catalogs to market due to the extremely high costs of captioning and localization,” Huidor said.

The traditional transcription process involves an individual transcriber listening to a piece of audio and manually converting every audio element they hear to text. It is clearly very labor intensive using trained specialists, and costly.

But it does produce accurate results.

AI transcription eliminates the need for a human transcriber and relies instead on automatic speech recognition technology. ASR uses language and learning models to interpret human speech and convert specific sounds (or phonemes) as written language.

Some of the most popular speech-to-text software is provided by Google, Azure, IBM, and Dragon Professional.

The upside of using automated transcription is the ability for companies to scale more of their output, to keep pace with huge global demand and to slash the costs of the whole exercise.

The main downsides, as outlined by Vitac, are inaccuracy. AI system tend to deliver poor quality results when the input recording is poor, when there are more than one speaker and when the audio contains a substantial amount of overlapping speech. Other factors that can inhibit the AI’s ability are when speakers have diverse accents or dialects.

“All these variables can substantially impact AI’s ability to interpret and represent the audio of a recording and result in a final transcript containing a substantial number of errors,” Vitac says.

Its prescription to achieve “exceptionally high rates of accuracy” is to match automation with human experts. Not coincidentally this is exactly the service it offers.

Broadcasters and publishers are a little reticent to rely on AI transcription given that tools to date have not proved fool proof. The BBC, for instance, values the trust that viewers put in the veracity of its output more than most broadcasters. It also faces increasing pressure to cut costs. It is exploring and evaluating AI tools which is a route that it advises others to follow.

 

Vanessa Lecomte, localization operations manager at BBC Studios, telling language information site Slator that for all the benefits that AI has in localization, it “must match BBC’s quality standards at a minimum.”

She said, “The main question is whether AI can improve current processes, increase speed to market, and reduce costs.”

Lecomte advised balancing opportunities against the risk. “These technologies offer the potential to speed up the process, which in turn enables you to localize more content, reach new markets, but it shouldn’t be done to the detriment of quality or of a well-respected industry. So do the right thing and commit to a thoughtful localization strategy.”

The BBC is also addressing AI in dubbing using synthetic voices. Lecomte described the current dubbing process as “time-consuming and expensive involving many technical and creative talents.” She said her division is exploring the capabilities of AI dubbing technology to try and deliver more content, faster, and still meet quality standards, adding that this should be done acting responsibly in regards to talent rights.

Anton Dvorkovich, CEO & Founder of Dubformer, also flagged the industry responsibility of establishing regulations around the ethical use of human voices.

He also believes AI dubbing is “poised to dramatically transform the media industry…with solutions that cut production costs by 30-50%.

“For now, investors and the media are struggling with the challenge of evaluating new solutions. However, the focus is shifting to the potential costs of emerging tools and their impact on the media industry,” he wrote in an op-ed for Streaming Media.

Solutions range from those like Papercup and Deepdub where humans finalize the AI-powered dubbing to “DIY translation tools” aimed at enabling freelance content creators to translate their videos with AI. One such solution, from Heygen, relies on natural-sounding speech synthesis and text-to-speech software developed by Eleven Labs.

He predicts that the introduction of an “AI Dubbing Manager,” or proof listener, tasked with fine-tuning AI dubbing systems or types of content. This role could include listening to the automatic voice overs to grasp cultural nuances, refine voice modulation, and make corrections. Some actors and interpreters may transition into this profession as it evolves, he suggested.

There could be Creative Directors for AI-enhanced productions to guide creative content developed through AI dubbing while the market for actors to license their AI-generated voices will grow. “More tools will enter the market, enabling individuals to generate their voices with AI. Actors will be able to create new voices based on their own.”

AI-Powered Localization and Captioning Tools

Software developer Enco introduced AITrack and ENCO-GPT, which both use ChatGPT to generate language responses from text-based queries for automated TV and radio production workflows.

AITrack, for instance, integrates with Enco’s DAD radio automation system to generate and insert voice tracks between songs. It leverages synthetic voice engines to produce natural-sounding, engaging content between songs.

ENCO-GPT could be used to condense a lengthy written news article into a few sentences, or inject breaking news updates within live ad breaks or automatically creates ad copy on behalf of sponsors.

Company president Ken Frommert sees an opportunity to go bigger with both solutions. “We see opportunities to convert a morning or afternoon drive radio show into a short-form podcast, or summarize an 11:00 p.m. local news program for the TV station’s website…. It offers a seamless way to publish content in diverse forms.”

LEXI Recorded, a VOD automated captioning solution from Australian firm AI Media, claims 98% accuracy, “comparable to human captioning,” and even higher with the use of custom dictionaries or topic models. Its use is priced from 20 cents per minute.

“We are not just meeting but exceeding the demands for high-volume, quick, and precise captioning of recorded content,” said AI-Media’s Chief Product Officer, Bill McLaughlin who will present the product at NAB Show in April.

Captions offers an AI-based video editing app and a solution for automatically generating subtitles. Both products are aimed at content creators and marketers.

It also offers an in-house voice cloning tool trained on licensed audio recordings to translate users’ audio into 28 other languages or use an AI voiceover to narrate the content from scratch.

Gaurav Misra, CEO and cofounder says Captions’ approach to video editing software is different because its tools are designed for specifically editing talking videos. “Most video production editing is focused more on aesthetics like filters and colors, whereas our focus became more about conveying an idea or experience,” he told Rashi Shrivastava at Forbes.

Vitac’s claims its own AI captioning solution, Verbit Captivate, stands apart from “generic” ASR engines in being designed, developed and built, inhouse. “Whereas other AI captioning vendors either provide an engine or a service, Vitac is unique in that we own both. And because of that, we can change, update, upgrade, and customize customer offerings, tuning our solutions to individual customer needs, creating an offering that achieves accuracy and results on a personal level.”

Additionally, it pairs the tech with “human backup” — specialists who boost performance with prep, pre- and post-session research, and live-session monitoring.

Cineverse’s MatchCaption, targets bulk film, television and video libraries localization “at significant scale.” It claims its generated captions are “perfectly timed and formatted according to industry standards, then auto converted into multiple caption/subtitle formats, to meet the specifications of all streaming platforms.

It also claims its system can complete the same tasks which currently cost content owners $60-100 for less than $10 per program hour, “and a full feature film can be completed, and quality checked in less than one hour — an 85% reduction in cost and 90% reduction in time.

 


The Ambition to Make Movies More Accessible

NAB

Captioning on-screen content is fast becoming the standard on TV but many in deaf and hard of hearing communities continue to call on all film screenings to include permanent, burned-in open captions.

article here

In a Los Angeles Times article, “Captions took over TV. Why can’t they win the silver screen?“ Sonja Sharp states that despite roughly half of all TV viewers use captions most or all of the time regardless of whether they are hearing impaired, captioning at theatrical screenings in the US is patchwork at best.

The term “open captions” refers to dialogue and audio description projected like subtitles on the screen. All theaters must provide subtitles in accordance with the Americans with Disabilities Act. Advocates however prefer open captions for film screenings due to their more inclusive user experience. Open captions are permanently fixed and timed to a film, they can’t be turned ‘off’ and the universality of their inclusion would remove what some see as the stigma of having to have their needs addressed as ‘special’ in cinemas.

Movie theaters are also required to provide and maintain closed captioning and audio description equipment for digital films that are produced with accessibility features.

“These devices convey captioning to the individual user on the false premise that the rest of the audience do not want to be ‘bothered’ by captioning,” Howard Rosenblum, CEO of the National Association of the Deaf, told Chase DiBenedetto at Mashable. “In other words, the deaf and hard of hearing patron has to endure discomfort of these goggles or contraptions for the presumed comfort of everyone else.”

The National Association of the Deaf (NAD) states that the two types of captioning equipment available in theaters are Sony Entertainment Access Glasses and CaptiView. Additionally, theaters are required to provide notice to the public about the availability of accessibility features and ensure that staff is available to assist patrons with equipment.

But getting open captions in theaters is far more complicated than it should be both from a logistical and technical point of view, advocates say.

“You’re looking up and down and up and down and up and down through the whole movie,” deaf moviegoer Dani Duran told the Los Angeles Times. “If the captioning goes off in the middle of the movie, it’s ‘sorry, too bad, and we’ve missed 20 or 30 minutes of the movie’ before it can be fixed.”

In 2023, Variety reported that Sundance Film Festival jurors Marles Matlin, Jeremy O. Harris, and Eliza Hittman walked out of a film screening after Matlin’s closed captioning device malfunctioned and no other captioning alternatives were available to her and other deaf and hard of hearing audience members. 

Filmmaker Alison O’Daniel, who is deaf/hard of hearing, wrote in Variety, “I am always hesitant to complain about accessibility, but I am hyper aware of what disabled people have to deal with to gain that access. I carry a CaptiView into a theater and feel people look over as I adjust it. I leave screenings with a headache from looking back and forth between the device and the screen.

She added, “I look forward to a time when captions are such an obvious benefit that films without captions are a part of cinema’s past, much like silent films.”

Filmmakers generally want to be accessible to the widest audience possible but it seems they are facing technological and procedural constraints. In the meantime, closed captions remain a way for filmmakers and theaters to provide a compliant solution without taking a “visible stand” on the issue, according to Matt Lauterbach, a filmmaker and accessibility advocate who founded All Senses Go.

Lauterbach explained to 3Playmedia that captioning technology can be cognitively draining, straining on the eyes, and even cause users to miss content in screenings due to the need to look back and forth from a device to the screen.

“It’s a tough user experience,” he said. “The device needs to be set to the proper theatre. You might get a caption device set to theater 7, and it’s set to theater 6. You then need to bring it back to get it fixed [during the movie].”

On top of incorrect theater settings, dead batteries and uncharged devices are a common issue, not to mention theater and festival staff who aren’t trained on how to use or troubleshoot captioning devices.

There are open-caption screenings but they are far from universal and you have to know where to look.

“You can’t schedule the open-caption slot at a lousy time, not promote it and then complain community members aren’t coming,” Melissa Greenlee, founder of Deaf Friendly Consulting, said to the Los Angeles Times. “To not give us options during prime-time weekends … is like saying: ‘You don’t matter.’”

ASL is even rare in theaters and on streaming services.

The cost of creating an open-captioned print is cited as a barrier, but Lauterbach said that the Digital Cinema Package (DCP) can actually be formatted as both closed and open captions without a need for additional quality control or much of a difference in overall cost. When a captioner creates a DCP caption file, it’s a matter of toggling settings on and off via the DCP.

In addition, because open captions are part of a video, they are supported by all video players and devices. Open captions eliminate rendering inconsistencies across different video players and devices.

Many accessibility advocates say that the cost of not including a major group of people is greater than the cost of adding open captions or subtitles to film screenings because of the enormous segment of consumers being excluded, Salon’s Alison Stine reports.

United States Deaf, hard of hearing, and hearing loss communities consist of more than 30 million people. Plus, millions of non-native English speakers, neurodivergent audiences, and viewers who prefer media with captions turned on make up additional viewing groups who have helped fuel the unprecedented usage of captions in recent years.

“Ask deaf and hard of hearing people who use closed captioning at movie theaters and you will get many stories from them about malfunctions, battery problems, disconnects, missing dialogue lines, mix-ups of captioning from the wrong movie, limited quantities of the devices, physical discomfort with goggles, difficulties keeping the cupholder contraption in the line of sight, staff errors, and much more,” Rosenblum explained to Mashable. “By contrast, open captioning provides deaf and hard of hearing people with truly equal access in that they can go into the theater and watch a movie without any extra effort or having to secure any equipment.”

 


OpenAI’s Sora: It’s the Beginning or the End of Video and Either Way It’s a Big Deal

NAB 

OpenAI seems to delight in pulling rabbits from a hat and was more than aware of what its latest research project would do when it alerted the internet. Everyone’s gone wild for Sora, a new diffusion model being tested which can generate one minute video clips from just a single text input. To prove what it can do OpenAI dropped some videos online generated by Sora “without modification.” One clip highlighted a photorealistic woman walking down a rainy Tokyo street.

article here

“Every single one of [them] is AI-generated, and if this doesn’t concern you at least a little bit, nothing will,” tweeted YouTube tech journalist Marques Brownlee. “This is simultaneously really impressive and really frightening at the same time,” he added on his YouTube channel.

A blog post on the website of nonlinear editing software Lightworks declared, “Sora’s almost magical powers represents yet another seismic shift in the possibilities of content creation.”

“It’s incredible and scary” says Erik Naso of Newsshooter.

“Sora is a glimpse into a future where the lines between creation, imagination, and AI blur into something truly extraordinary,” gushed Conor Jewiss at Stuff.

Benj Edwards of Ars Technica thinks OpenAI is on track to deliver a “cultural singularity” — the moment when truth and fiction in media become indistinguishable.

“Technology like Sora pulls the rug out from under that kind of media frame of reference. Very soon, every photorealistic video you see online could be 100 percent false in every way. Moreover, every historical video you see could also be false.”

What has excited the AI and artistic community so much is the cinematic photorealism of the videos produced by OpenAI’s algorithm which seems “to understand how things like reflections, and textures, and materials, and physics, all interact with each other over time,” said Brownlee.

In its research paper Open AI states the model deeply understands language, enabling it to accurately interpret prompts and generate compelling characters that express vibrant emotions.

Sora can also create multiple shots within a single generated video that accurately persist characters and visual style.

OpenAI further states it is teaching the AI to understand and simulate the physical world in motion, with the goal of training models that help people solve problems that require real-world interaction.

Two videos in particular grabbed attention. “This is one of the most convincing AI generated videos I’ve ever seen, says Brownlee of a video made with this text prompt: “A movie trailer featuring the adventures of the 30 year old space man wearing a red wool knitted motorcycle helmet, blue sky, salt desert, cinematic style, shot on 35mm film, vivid colors.”

“This looks like it could be an actual film trailer,” says Theoretically Media’s Tim Simmonds.” I mean that there’s nothing really in here to majorly indicate that this is AI generated.”

The other, featuring an aerial flyover, was spun-up from the prompt: “Historical footage of California during the gold rush.”

“The drone footage of an old California mining town looks really, really pretty great,” Simmonds says. “And even as the camera makes this turn here, the buildings stay intact, they don’t start to shift and warp and morph into weird things.”

Brownlee thinks it demonstrates “all sorts of implications for the drone pilot that no longer needs to be hired, and all the photographers and videographers whose footage no longer needs to be licensed to show up in the ad that’s being made,” he says.

“It’s also very capable of historical themed footage,” he adds. “This is supposed to be California during the gold rush. It’s AI generated but it could totally pass for the opening scene in an old western.

Which begs the inevitable question, How long until an entire ad with every single shot is completely generated with AI? Or an entire YouTube video, or an entire movie?

Simmonds still thinks we are a way out from that “because [Sora] still has flaws and there’s no sound [no audio/dialogue sync] and there’s a long way to go with the prompt engineering to iron these things out,” he says.

Naso agrees that Sora “could change the game for stock footage,” adding that the next stage for AI prompt filmmaking is dialogue-based scenes. “So far, these examples are more like b-roll.”

Nonetheless, even at the pace of AI development it seems OpenAI has caught everyone napping.

Rachel Tobac, a member of the technical advisory council of the Cybersecurity and Infrastructure Security Agency (CISA), posted on X (formerly known as Twitter) that “we need to discuss the risks” of the AI model.

“My biggest concern is how this content could be used to trick, manipulate, phish, and confuse the general public,” she said.

OpenAI also says it is aware of defamation or misinformation problems arising from this technology and plans to apply the same content filters to Sora as the company does to DALL-E 3 that prevent “extreme violence, sexual content, hateful imagery, celebrity likeness, or the IP of others,” as Aminu Abdullahi reports at TechRepublic.

Others flagged concerns about copyright and privacy, with Ed Newton-Rex, CEO of non-profit AI certification company Fairly Trained, maintaining: “You simply cannot argue that these models don’t or won’t compete with the content they’re trained on, and the human creators behind that content.”

Anticipating these concerns, OpenAI plans to watermark content created with Sora with C2PA metadata. However, OpenAI doesn’t currently have anything in place to prevent users of its other image generator, DALLE-3, from removing metadata.

OpenAI said it is engaging with artists, policymakers and others to ensure safety before releasing the new tool to the public. However, its get-out clause is that despite extensive research and testing, “we cannot predict all of the beneficial ways people will use our technology, nor all the ways people will abuse it.”

The Microsoft-backed company is valued at $80 billion after a recent injection of VC funds. “It will become impossible for humans to detect AI-generated content by human beings,” Gartner analyst Arun Chandrasekaran warned TechRepublic. “VCs are making investments in startups building deepfake detection tools, however, there is a need for public-private partnerships to identify, often at the point of creation, machine-generated content.”

Sora joins a chorus of other text to video generators such as Runway and Fliki, the Meta Make A Video generator, and the yet-to-be-released Google Lumiere.

Question: Has Apple taken its eye off the ball? Answer: Maybe not. Its researchers have just published paper about Keyframer, a design tool for animating static images with natural language.

As Emilia David at The Verge points out, Keyframer is one of several generative AI innovations that Apple has announced in recent months. In December, the company introduced Human Gaussian Splats (HUGS), which can create animation-ready human avatars from video clips. Apple also released MGIE, an AI model that can edit images using text-based descriptions.


Profit quest drives content licensing deals between rival SVODs

StreamingTV Insider

The pursuit of streaming profitability over subscriber numbers has seen the major studio groups turn their attention to content licensing deals. The sweet spot is finding the balance between exclusivity and non-exclusivity, and in the case of non-exclusivity using that content to drive licensing revenue.

article here

New research published by Ampere Analysis shows that after four years of major studios employing a walled-garden approach to the distribution of their TV content on streaming, licensing is steadily making a comeback.

Between 2019 and 2021, media company studios s began to pull content from Amazon and Netflix and other licensing partners prior to debuting their own SVODs, ensuring that content exclusivity was a major part of their launch proposition. A notable move was WarnerBros’ removal of Friends from Netflix ahead of HBO Max launch.

“Having content exclusively available particularly when its hugely popular or part of major franchises can help when building a brand at launch and therefore helps in penetrating an increasingly competitive market,” said Rahul Patel, research manager at Ampere Analysis.

Now the industry is moving to a new period where cross licensing is becoming a lot more common. “This also reflects the state of maturity these platforms are reaching,” Patel told StreamTV Insider. “They now better understand what content is keeping consumers engaged or attracting new subscribers and what other content can be additionally monetized through cross licensing.”

Ampere’s analysis highlighted significant recent increases in catalog overlaps between different platforms. Netflix and Warner Bros. Discovery, for example, started 2023 with 22 TV seasons shared between them and ended the year on 68 (including HBO/Max shows Band of BrothersInsecureBallers).

Amazon Prime Video and Peacock’s season overlap jumped from 156 to 275 in the same period (including MonChicago Fire and Chicago P.D).

Its data also illustrates the scale of popular content that has been retained by the studios “highlighting the potential and room for growth in licensing.”

Disney has most muscle here with 148 TV shows identified as having licensing power still only available on its platforms at end of 2023. That is about to change when a major Netflix deal for around 14 Disney and Fox series rolls out on Netflix between now and January 2025, including Prison Break, and How I Met Your Mother. This totals in excess of 90 TV seasons.

For context, Ampere data shows Netflix and Disney+/Hulu had an overlap of 152 TV seasons in Decemeber2023, up just 13% from January 2023’s 134. So the deal in its entirety represents a big change.

Paramount has the second largest cache of high value content. Star Trek and its spin-offs are core to the Paramount brand and are understandably retained exclusively for Paramount+ but these “also reflect some of the most popular licensing assets if they choose to go down that route,” said Patel.


AVOD platforms are also licensing partners. NBCUniversal has been most active in this area, sharing the most content not just between SVOD platforms but also with AVOD platforms. WBD is distributing over 2000 hours of its content, including through Warner Bros-branded FAST channels on Tubi and The Roku Channel.

Ampere’s research also shows that studios are considering changing their strategy around core IP exclusivity. WBD for instance has been very protective of its theatrical slate, in contrast to Paramount and Universal whose theatrical releases have often landed on other platforms a few months after the first pay window, noted Patel. “Yet in Q3 last year WBD licensed a whole host of its movies to Amazon including DC titles and then rolled them out to Netflix and Tubi for limited periods.”

Reaching wider audiences through licensing

While the principal driver to cross license is for increased revenue, another benefit for studios is to reach wider audiences with their content.

“I see licensing as part of a marketing tool,” Patel said. “It’s not by accident that Warner Bros. licensed a set of DC titles to other platforms like Amazon and Netflix on the eve of releasing Aquaman 2 at the end of last year.”

WBD is also featuring its Oscar winning hit Dune prominently on Netflix in the run up to the sequel’s release next month.

“By licensing these franchise titles to platforms with a large audience, studios can market to more people and hope for a positive return as more people engage with content in its original home.”


AMC+, a far smaller streamer, sought the wider reach of Max when it created a ‘streaming pop-up’ on Max to carry seven AMC series for two months beginning September, including Anne Rice’s Interview with the Vampire and Killing Eve.

WBD itself will be hoping to engage new audiences in seeing Sex and the City revival series, And Just Like That, on Max after making all six seasons of HBO’s original series available from April on Netflix.

Similar activity has not yet been witnessed by Apple TV+ which seems to be pursuing a policy of quality of quantity when it comes to content.

Ampere’s analyst doesn’t rule out a change in Apple’s strategy, particularly given the shifting chairs of bundles, mergers and acquisitions in the industry.

Amazon, meanwhile, has licensed some unscripted originals to WBD for distribution on broadcast TV in some EU markets in what Patel called “another change in the playbook for streaming originals.

“Netflix meanwhile is at the forefront of the streaming game and perhaps calculates less of a need to license out content,” he said. “It has found other avenues to generate revenue such cracking down on password sharing and the launch of its ad tier.”

The industry can expect more licensing deals for high profile titles to be struck in 2024 between major VOD providers. Patel added that cross licensing can be seen “as a first step in wider partnerships between platforms” as the industry reshapes itself with bundles, mergers and acquisitions.

Broadcaster, SVOD overlap ahead for the European market

Ampere’s research is based on the U.S SVOD market and the conditions around exclusivity around what the studio and streamers catalog looked like at the end of 2023.

“The walled garden approach is slightly differently in Europe because some of the major U.S streamers aren’t operating in certain markets,” said Patel. Peacock, Hulu, Max for example have not launched in the UK. Ampere will publish research in the market later this year.

However, Ampere predicts more licensing between UK broadcasters and streamers following that between Disney and UK commercial broadcaster Channel 4 which saw ten Disney series including Abbott Elementary and X Files made available to the broadcaster’s eponymous streaming platform.

ITV continues to strike carriage deals with Warner Bros for TV and on streaming platform ITVX including the Harry Potter franchise and The Vampire Diaries.

Patel said, “We expect the partnership and overlap between broadcaster led VOD and SVOD platforms both in the UK and other major European markets to increase, in the same way we are seeing cross licensing between major SVOD providers in the U.S.”