This right here, or at least the thought of this, is why in the not so far future, businesses can't (won't?) be using these LLMs services.
You cannot risk companies like Anthropic, OpenAI or their business partners like Microsoft having unfettered access to proprietary data on your company/businesses.
It it likely that they or rogue employees will use the information to make a profit? It's pure speculation, but I'd say more than likely, and we will never hear about it or read it on the news unless there's whistleblowers in high enough positions to know about it.
Assuming you and your employees aren't careful with what data you share, they will have intimate knowledge about your company from files and conversations logs. Likely personal user data too which they'll gladly create databases to link to and create extensive profiles on you, your employees and your businesses.
It's not far-fetched to see them leveraging insider information shared with LLMs to play the stock market, leveraging data against competing businesses in other markets they might want to explore, and likely a bunch of other things that are escaping me right now as I write this.
At the end of the day it's on those people for sharing such sensitive data, but it's not like these AI companies are innocent and won't gladly exploit every little byte of data without telling you, we know it happens.
Its not "in the future", its already here, right now. Sensitive IP gets processed on rented or local GPUs running open weight models. Sometimes its required due to data privacy laws, sometimes its because the people running those shops are not trusting OpenAI/Anthropic ZDR policies being observed. Normal, boring, US and EU-domiciled enterprises do this on a regular basis, and there is an industry of companies helping said boring companies set things up.
It is even worse. Just them knowing or having an inkling that someone is close to achieving a solution might be enough for them to spawn a team of thousands of agents and outrun you.
> This right here, or at least the thought of this, is why in the not so far future, businesses can't (won't?) be using these LLMs services.
Or they can simply turn off the setting that allows OpenAI to use their chats to improve the model. Or they can use the API where it is off by default.. If the advantage of using OpenAI models will be significant enough for the business in question, these are the options.
>You cannot risk companies like Anthropic, OpenAI or their business partners like Microsoft having unfettered access to proprietary data on your company/businesses.
You do realize that businesses that have contracts with them can specify if they want their data used for training or not right?
One example to illustrate: when talking about dying, Americans would kick the bucket, Poles would kick the calendar, French would break the pipe, Germans would drop the spoon, and the Spanish would pull the leather. But they all would be referring to the US's trust among its allies.
In the US "Could care less" is the equivalent to the British expression "Couldn't care less". The intention and meaning is the same. So the original comment did mean to say what they said.
My point was that what may look like a mistake or a contradiction is just cultural development of natural language.
Another example is the British expression "It would be cheap at half the price"
maybe it's regional? I always thought "Could care less" was in the same class as "irregardless" or "daylight saving time" -- in common use yet technically incorrect
I think you misunderstood the emphasis there - yes Europeans love football.
The point is that it's so blatant for such little gain, it's still a game. I.e. if he is willing to do that so openly for such low stakes, imagine what he does behind closed doors when it matters.
Most have been horrible or horribly mediocre. Financial success doesn't necessarily mean something is good either. Also, some of these are more like: "very loosely based on the video game, so much so that you have to squint very hard to see it." Stuff like Detective Pikachu, or Minecraft are just aesthetics and have little to do with the actual games.
Jun. 10, 2016 Warcraft Film
Dec. 21, 2016 Assassin's Creed Film
Jan. 27, 2017
Resident Evil: The Final Chapter
Film
Mar. 16, 2018 Tomb Raider Film
Apr. 13, 2018 Rampage Film
May 10, 2019 Pokémon Detective Pikachu Film
Feb. 14, 2020 Sonic the Hedgehog Film
Dec. 4, 2020
Monster Hunter
Film
Apr. 23, 2021 Mortal Kombat Film
Jun. 25, 2021 Werewolves Within Film
Feb. 18, 2022 Uncharted Film
Feb. 24, 2022 Halo TV series
Apr. 5, 2023
The Super Mario Bros. Movie
Film
Apr. 8, 2022 Sonic the Hedgehog 2 Film
Jul. 14, 2022 Resident Evil TV series
Jan. 15, 2023 The Last of Us TV series
Jul. 27, 2023 Twisted Metal TV series
Aug. 25, 2023 Gran Turismo Film
Oct. 27, 2023 Five Nights at Freddy's Film
Apr. 5, 2024 Knuckles TV series
Apr. 10, 2024 Fallout TV series
Aug. 9, 2024 Borderlands Film
Dec. 20, 2024 Sonic the Hedgehog 3 Film
Apr. 4, 2025 A Minecraft Movie Film
Apr. 25, 2025 Until Dawn Film
Dec. 5, 2025 Five Nights at Freddy's 2 Film
Jan. 23, 2026 Return to Silent Hill Film
Jan. 30, 2026 Iron Lung Film
May 8, 2026 Mortal Kombat II Film
PS: Note the question being asked is "Will it be any good?" Not "Will it be profitable?"
Yeah, most of the list is pretty dismal, but there are some bright spots. Sonic movies. Fallout, Last of Us (and don’t forget Cyberpunk: Edgerunners) for TV.
The original Castlevania animated show was good, shame about the Warren Ellis situation. Would have loved to see him write Nocturne as well (which was fun but just had a much weaker script).
Please switch off Fox News and fetch real world updates. Europe's dependence and weapon system fragmentation is by USA design; Europe has allowed the US to have this advantage as part of the security architecture, alongside other goodwill handed to the US. Europe is rapidly rearming while the American mil. complex is throwing tantrums, but you can't have your cake and eat it too.
> At least 69 police officials have been accused, charged with or convicted of misusing Flock’s system or other license-plate readers for unauthorized purposes, The Post has found, and in at least 15 of these cases, someone outside of the police department — including victims, activists and journalists — first identified the potential misuse.
> Records and interviews show that departments across the country do not regularly audit their officers’ Flock use, a practice recommended by the company, and some law enforcement agencies have no formal training programs or written policies describing how the technology should be used.
One aspect that people seemingly aren't talking about is the impact this has in the model's creativity. Because the model will nudge each word towards group A vs group B, you're losing on creativity, especially more so if the nudge isn't a gentle 55% but something like 70% or 80%. So essentially they're forcing the model to be less creative for the upside that the longer the text the easier it is to detect the watermark.
This is even worse in such forms of writing like coding, where there's even less choices the model can make on what the next token should be. Plain text in code will obviously be watermarked, that includes comments. But the code itself might get watermarked by choosing certain code over others more often.
I'm inclined to believe the models will be instructed to not watermark code, especially since it's harder to detect reliably because the shorter the body of text the harder it is to detect, but who knows what Anthropic and all the other AI labs will decide to do in the future.
EDIT: Also for those in the comments who are naive enough to think Anthropic is doing this just because the EU said so and not because it's beneficial to them (and all other AI labs), well, you are indeed naive. Identifying code will be paramount in training future models because the more synthetic data you feed it, the more cannibalization happens, the worse the models will perform over time due to lack of good data, among other such reasons as selling AI detection services to colleges, and a plethora of other reasons.
As per the link, the words in the green and red groups are calculated dynamically, so it's not like the model is going to be told "use 'unique' over 'unusual'" and suddenly writing from the model will contain the word 'unique' far more often than 'unusual'. So I'm not sure it's clear that this has an impact on creativity as such?
That being said, I do question how this will apply to code as opposed to prose. Even data dense text (ie, if you ask Claude to evaluate what running shoe to buy, and it spits back a list of options with reviews and prices) may struggle.
What it probably will work well at it flagging the current tsunami of entirely AI generated novels on Amazon/Kindle, which is...honestly not without value.
> Identifying code will be paramount in training future models
True, but note that this strictly allows providers to identify text generated by their own models. If Anthropic wants to filter out GPT generated text in their training data, they'll need to feed it through an OpenAI API, which is implausible. So it might help on the margins, but I don't think it solves the problem of model collapse.
The model is not given a list of words. The model has a seed (key) which at inference splits all possible next token into list A and B, and nudges (has a bias) for list A. How many possible options the next word has depends on what came before, some will have a few hundred, some will have hundreds of thousands of possibilities.
The model essentially is just doing what it always does which is predict the next token, but the next token now is split into two and nudged towards one side more often than the other which is how over a body of text identifies if the text was in-fact AI or not. Which is also why the shorter the text the harder it is to identify.
As for the second half, I predict that almost all American, Japanese and European labs will use watermarking at some point (as well as SynthID for other generative AI). It doesn't matter if they all have unique internal seeds because they'll give people the tools to ID AI, be it by selling it (unlikely for most use cases , but likely gonna happen for academic where they'll provide some value app that bulk checks student works), or more likely make it free to check like SynthID where you simply ask Gemini if the image has SynthID. At which point they can just pay each or use each other's tools to check.
The wildcard are the Chinese models, but those will likely force some sort of watermarking as well, if not for the global market, for the CCP's benefit.
The details will depend on the exact implemention, but I don't think this necessarily effects output quality.
The models "natural" output is the result of a series of random numbers. The watermark works by biassing that series towards a different series of numbers. Assuming that second series is cryptographically secure psuedo-random, the even distinguishing the biased sequence from true random would be impossible with compromising the key or prng.
As an extreme, suppose your prompt was public, and the model seeded its PRNG with a secret key instead of a genuine random seed. Such an output is not meaningfully different from one based on a true RNG, but can be trivially fingerprinted by someone who knows the keys.
In practice, I am doubtful they have a scheme that is both practically useful and cryptographically secure. However, there is a lot of room below cryptographically secure that is still just as good for all other purposes.
It is worse than that, it can make text unintelligible. I have Claude explain what is happening in a PR, and the technobabble and use of rare words have to look up in a dictionary make it difficult to understand. I need another LLM to translate what the LLM is saying.
And I pray it is because of the EU AI Act, because it is worth giving them my personal information to prove I am not in the EU so they can turn this off. Heck, I will pay more to avoid this crap.
If it were not so noticable, I would shrug it off. But it has made things clearly worse this year.
This is purely my opinion. I see 70mm as mostly a movie scam to extort exorbitant amounts of cash from a ticket for something that's vaguely noticeable outside the aspect ratio itself.
Note that I'm not disputing that there is a visual difference, there is, the image is clearer, shaper, with better darks, but the entry fee paired with how out of the way 70mm is, it's just ends up feeling like FOMO.
My opinion doesn't hold much value since I just dislike the cinema experience in general. I like being able to control the environment I'm in when watching something (pausing, eating, loudness...) and whenever I go to cinemas I have to deal with people having their smartphones screens brighten up, people yapping and laughing at things that aren't funny, eating noises, etc...
I don't think the extra production cost, for the Odyssey at least, is likely to be offset by extra box office income. If you read about it, the massive cameras and limited running time of one load of film (less than five minutes) sounds very burdensome. That's not to mention the obscene cost of the film stock, as well as developing dailies and having to edit from film (they actually cut and pasted film together for the Final Cut of the IMAX version).
The trailers also looked much worse than the film which I suspect is related to difficulties with this whole process.
I am with you. I completely believe that the difference is there and that it is probably more immersive.
However, these discussions always feel similar to MP3 192kbit/s vs AAC 256kbit/s vs lossless vs lossless 24-bit 192KHz. Trained listeners can hear the difference between MP3 192kbit and AAC 256 kbit. Perhaps even AAC 256kbit vs lossless. But IRL it does not matter:
- Some people listen on very crappy headphones/speakers.
- Some people listen on good headphones with AAC, but even with AAC on, external sound leakage probably has a larger effect on the perceived signal than compression artifacts of AAC 256kbit/s.
- Most of the time people listen to music in the background, so there is not a lot of attention on it anyway.
- Many (most?) people primarily listen to melody and lyrics and not to the fine grained details of how the instruments sound.
Similarly, a lot of people who watch movies, as long as the image/audio quality is good enough, they focus on the storyline, the dialogs, the general atmosphere, etc.
I am not denying that there is a market for audiophile audio or the best quality image for movie visuals buffs, but for most people these are just overpriced gimmicks. I mean, we used to watch movies at home on crappy, small CRT TVs, and many of us still have fond memories of many movies despite seeing it in a format that could as well be from the stone age compared to even a standard projector used in an average cinema nowadays.
To avoid any straw manning: I am certainly in favor of all the improvements in tech we have seen, it's more that the experience does not linearly with it. More like exponential saturation.
You cannot risk companies like Anthropic, OpenAI or their business partners like Microsoft having unfettered access to proprietary data on your company/businesses.
It it likely that they or rogue employees will use the information to make a profit? It's pure speculation, but I'd say more than likely, and we will never hear about it or read it on the news unless there's whistleblowers in high enough positions to know about it.
Assuming you and your employees aren't careful with what data you share, they will have intimate knowledge about your company from files and conversations logs. Likely personal user data too which they'll gladly create databases to link to and create extensive profiles on you, your employees and your businesses.
It's not far-fetched to see them leveraging insider information shared with LLMs to play the stock market, leveraging data against competing businesses in other markets they might want to explore, and likely a bunch of other things that are escaping me right now as I write this.
At the end of the day it's on those people for sharing such sensitive data, but it's not like these AI companies are innocent and won't gladly exploit every little byte of data without telling you, we know it happens.
reply