Most experimental physics and other natural sciences are strongly driven by their theoretical siblings, i.e. in particle research nothing gets built without a solid theoretical foundation of what you expect to find (or where you expect existing theories to break down), the same is true in other areas, no one is doing an experiment in quantum physics before they have a solid theoretical understanding of the effects they try to see. I think AI can come up with great experiments. And if epxeriments lead to results that are unexpected AI can help with that as well.
So I'm greatly excited what AI will bring about in physics, more so than in math, because in physics it's clear that our fundamental theories are missing a big piece of the picture, and given how easily AI crunches through Millenium prize problems I think it's possible that AI will come up with a viable grand unified theory uniting quantum mechanics and gravitation, or produce new predictions in other areas. There's enough contradictory or unexplained observational data available to make a ton of progress on the theory side I think. Exciting times ahead!
Agreed! Many people are saying AI isn't really intelligent yet because it can't come up with genuinely new things. Maybe finding a rigorous formulation of QFT / high-energy physics would be a great test for whether they are!
Isn't there anywhere to "go" from here? In the last decades, introducing new high level abstractions on top of existing paradigms naturally had everyone move up the ladder and work at the next higher level, why should this be different these days? Do we think AI will reach the top of the abstraction ceiling, so there's no where to go from here?
This isn't abstraction though. Outsourcing is a better term.
If things continue moving up that latter, you will see that your agent/agency will pass the buck too. But there should always be some last turtle. Maybe that turtle will be the human that thought he was climbing the latter, who knows.
Coding via LLM is not similar to using an abstraction. Imagine a car. The controls like steering wheel, the pedals, the gear levers. Those are abstractions.
But using LLMs are like driving using a remote control that has probabilistic behavior. You just loss what it feels to be in a car and you fail to improve as a driver because of the erratic remote control.
To be fair AI as we have it today wouldn't exist if everything was copyrighted and copyright was enforced very strictly across the web. Not sure if it's fair use and I get that it's not great if these companies try to monopolize knowledge (if they even can do that) but overall being more lax on copyright for fair use is good. Especially scientific knowledge should be openly accessibly to anyone, imagine what we could achieve if we could train LLMs on everything and make it accessible to everyone.
Damn, big security fuckup by Dropbox, how can they portray that as an issue with Lenovo's e-mail verification process? You should never allow linking of an existing account with a new login method without first confirming that the user is able to sign in with an existing method first! Everyone knows this allows easy account takeovers otherwise, that's such a trivial attack vector, truly a scenario you could pose to a junior security engineer in an interview.
I mean it's mostly about energy efficiency, a well insulated house is as easy to cool in summer as it is to heat in winter. Most heat pumps can also cool and new construction in most European countries like Germany is equipped with floor heating and central ventilation which in summer allows you to run around 18 degree Celsius hot water through the system to cool it, that yields around 20-40 Watt of cooling power per square meter, and that's often sufficient to keep things chilly. Mind you not American style 18 degree Celsius kind of chilly, but 20-25 degrees easily, which I think is quite nice. The bonus of this system is that it's mostly "radiative" cooling as the cold surface acts as an absorber of infrared radiation, I find it much more pleasant than having cold air blown into your face.
We equipped our founding era building with wall heating and will test out the heatpump based cooling next summer, if it's not enough I would consider installing a small AC unit, that's super easy as well.
Don't know what the article wants to say, there's no ban on AC anywhere in Europe and small units are quite affordable, you can get a single split unit from Mitsubishi for 500 € from Italy. The only nuisance is that installation is heavily regulated in Germany so they rip you off with insane cost. That's definitely a thing that needs to change in my opinion.
Seems hard to believe, I mean I'm not an AI absolutist but as a programmer it allows me to automate simple things like writing small patches or config changes in a matter of minutes where before I would have invested hours, so hard to believe that it won't at least bump productivity up by 20-60 % overall, it's also tremendously better at finding relevant content and curating it than me so that's another time saver unrelated to programming but relevant for any type of knowledge work. Friends in research and medecine and law also tell me that AI has drastically accelerated their rate of processing and ingesting any type of information and helped them to automate aspects of their jobs. So not sure but negligible isn't the word I would use here.
> it allows me to automate simple things like writing small patches or config changes in a matter of minutes where before I would have invested hours
One of the problems here is that things like this only count as "increased productivity" in the economic sense if it results in more income at the end of the funnel. Are you actually making more product faster in all your new-found free time?
Very few orgs I've worked in have actually been limited by the speed of programmers writing code. Usually it's communication overhead and management processes eating up most of the bandwidth - and those tend to expand to absorb any additional bandwidth.
Yup. And in my experience, AI has made the communication and decision making overhead significantly worse thanks to 100x longer docs, code that noone understands and enormous pressure by executives to magically produce software that just works.
That's just because most people are bad at communication to start with.
Look at all the shit github README before llm. Most repos didn't bother because it was effort. those few that did, were because of autistic focus. Of those, very few were "great" user documentation.
now llms are here, the effort barrier is gone, and now we're inundated with shit README that provide no value to users nor respect their time.
Once you start applying SKILL.md that revolve around fixing HOW to communicate value to users in a way that respects their time, you'll see this story change in noticeable places.
I doubt it will change everywhere because it also requires that you give a shit about doing this in the first place, so those people Ed Zitron talks about... business idiots won't care, because they're ultimately nihlist.
> One of the problems here is that things like this only count as "increased productivity" in the economic sense if it results in more income at the end of the funnel. Are you actually making more product faster in all your new-found free time?
This strikes me as a problem with how we define productivity in this context. If I can do my job in 1 hour, when it took me 8 before, am I not more productive in a rate-based sense? Sure, I might only work 5 hours in a week instead of 40, so my net production is the same, but my rate of productivity was certainly much higher.
Wealth created is still the same, and that's what an economist means by "productive" - you produced X amount of wealth (usually measured in dollar amounts).
It's like "work" meaning one thing to a normal person, but something very specific and counterintuitive in Newtonian physics (if you carry a fifty-pound weight in a circle and stop at your starting point, you've done no work).
I suppose I can understand that perspective, but disagree with it? We should measure time returned to individuals as wealth, maybe even the highest form of wealth.
Everything is ultimately a chosen definition, nothing is "real". I'd argue the definition is antiquated. In the modern world, free time being returned to individuals is perhaps one of the greatest examples of wealth.
In the sense measured in economics, only if the product at the end of the day is sold for the same nominal amount of dollars as before. It doesn't matter how many features or lines of code or bugs or anything like this you produced - it matters how much revenue your work produced.
So, if the product isn't selling any more expensive even though you're adding ~8 times more features, or if you're still getting paid for the full 40 hours even though you only really work 5, then economic productivity hasn't actually increased.
Yup, I understand. And through a very narrow definition of economics it makes sense. I just don't see the utility in viewing it this way.
You've returned one of the most valuable resources (time) back to a human. In my mind there's simply nothing more valuable, and no sign of of efficiency more poignant than that.
It is actually a very relevant real-world definition - even though I agree with you on the intangible value of time.
The reason why it is very relevant is that, in the way we have unfortunately structured our economy, the amount of work you have to do to be able to eat and enjoy your time will always expand to grow to at least force you to maintain (if not increase) your economic productivity. That is, if AI enables engineers to deliver 10x more features in the same unit of time, then their employers will expect them to deliver 10x more features to continue earning their current pay.
So, while the AI did, in principle, allow you to just reduce your work schedule while getting the same pay, as you're providing the same results in 1/10th of the time, the economic system does not - your previous work is instead simply worth less than it used to be. Assuming you're not happy to take a 90% pay cut, you'll have to continue spending 40+ hour weeks, just using AI as well now.
It is really a deeper problem that "productivity" works in a factory producing widgets as output per labor-hour of widgets.
The Principles of Scientific Management was just wrong that these ideas extend universally but we are still doing this stupid Barnesian performance as if they do. It clearly makes no sense with knowledge work. It is stupid.
Then the efficient market hypothesis provides a type of circular proof because since the market is always right, if this didn't work we wouldn't be doing it. Another stupid idea we are not able to get past.
I mean I can deliver the same software in one month that I delivered in maybe 3 months before and for many things I don't need a dedicated frontend developer, designer or writer anymore. So for me it's the case!
For my job and basically everyone I know outside of programming, the productivity from AI is literally zero.
I think what this says is the productivity is in specific silos and why the experience of the utility of AI is so vastly different for different people.
Personally I feel a bit sad and annoyed sometimes, lots of folks I know that are working in some non technical roles e.g. lawyers in public departments or doctors in medecine come to me to show me their apps and tools and data science scripts they've written, and I have to admit their work looks totally respectable and they get it published on the app store. It can be a bit demoralizing as I feel like IT/software doesn't have any protected spaces for us engineers, whereas e.g. my friends in medecine or law feel fine as even though AI might come up with diagnoses or law expertise it won't be legally allowed to replace their function as there's a ton of red tape to ensue only certified lawyers and doctors can practice law or medecine. There's nothing like that in software engineering (at least in most fields) so in principle every "hack" can let their vibe coded data science results, software or script loose on the world. To a degree it feels quite disprespectful even. But it has always been like that I guess.
Exactly this. If those old-school professions didn't have protections they gained from lobbing for dozens of years, they would be facing the same reality we are.
If you think about it, a lawyer is an ideal profession to replace with AI. To a large extend doctors diagnosing based on inputs is the same. But they have the keys in form of their diplomas.
Not sure if I do anything wrong but Codex with Sol made such slow progress on my work, I iterated several days to do a redesign of my UI and it kept making small piecemeal changes then stopping and asking me for confirmation again and again even though I tried to get it to get bigger chunks done. Switched to Claude Code and had the redesign done in a short session within a day.
Maybe just the system prompt or my settings or the harness? Anyway, it seemed to me Codex was just deliberately being overcareful and wasting tons of tokens for a low-risk CSS / HTML refactoring with very minor breakage risk, while Claude got the job done immediately.
Would really like to know how their watermarking technique works, and if they can use it to store arbitrary information in the text (I assume they can if only in a limited way). I assume if the text is long enough they could add all kinds of metadata that would then be undetectable as the model that generated the text is the "cryptographic" key that encodes the data. I wonder if you can have another model run over the text and destroy the watermark. I predict an interesting cat and mouse game to develop.
The paper for it is open. The technique isn't really hiding information in the text itself, but by forcing some of the rolls to follow a specific pattern. LLMs work by estimating the most likely next token, so there's sometimes a list of possible candidates that would all work in the text (e.g. synonyms). At low "temperature", the output is a bit more deterministic and otherwise it's a weighted dice roll of which token/word to pick. An LLM can loop over existing text and figure out if the output matches something it would do, similar to checking chess moves against the best computed move for detecting cheating. But the LLM purposefully creates a pattern of alternating weighted rolls that are highly unlikely to appear in normal text, and that becomes the watermarking.
The upside is that this has very low false positive detection rate, but the downsides are many. It only works on longer pieces of text. The system is fragile, and small edits (or rewrites by a local model) can fool the detection. Only the owner of the model is able to re-run inference at this level, so data must be sent to them for evaluation. And sometimes the token output is basically 100% deterministic because the input asks for the straight answer to a fact, or to recite a quote verbatim. That leaves no room for watermarking at all, unless the model is able to lie.
A way to think of it is that any time the model faces a choice, it leaks some fractional bits of information. Instead of making the choice randomly, you can put information in those bits.
But unless you know the prompt, you don't fully know which choice the model faces. Surely a lot of coding space is wasted compensating for that uncertainty. Exotic prompts ("Use no more than five E's in four consecutive words anywhere in the text") seem like they will confuse the hell out of attempts to extract the bits from the output text alone.
> But unless you know the prompt, you don't fully know which choice the model faces. Surely a lot of coding space is wasted compensating for that uncertainty.
When you only need to encode one bit, the signal to noise ratio can be very low. If I try to write my own human words under the policy of "try a little bit to avoid the letter 'e' in every fifth word," then a sufficiently long text would still be 'watermarked' even if I only succeed in this dictum (e.g.) 10% more often than the baseline.
most people have at least two, but generally around four to six (or more, thanks internet) interactional styles (not "selves", just things which vary depending on context, subject, and people they are interacting with). probably more. some of it is due to simple physical comfort levels (right now I am in a physical position where capitalizing is more difficult); some might be due to talking to a peer group instead of a group of kids or a priest or boss at a job, etc. unsure how that will shake out with AI but it unnerves me.
There's more to it than encoding one bit. You also want to avoid false positives. You can encode a single bit by XORing all the bits in the UTF-8 encoding. But then you get a lot of incorrect hits. The lower your tolerance for false positives, the more it acts like you're actually requiring more bits in your payload.
The scheme that Scott Aaronson describes essentially uses a specific prng, and you can then check a certain function with relatively few tokens to get a sense of whether or not a model using that scheme generated the text
A few caveats: you need to know the key to the function (used when generating the text) and you need to know the bias it would introduce
The point is that you do not need to know the full prefix, just a modest sample set of contiguous tokens
In fact such ‘arbitrary’ constraints uniformly improve composition. Thus eg if I force a - largely arbitrary - technical glossary to be unrelentingly applied to a translation, every single sentence improves in quality.
> The system is fragile, and small edits (or rewrites by a local model) can fool the detection.
Well no, small edits wouldn’t fool the detection as long as the seeding only uses a small run of previous tokens.
And yeah full rewrites breaking it is by design. The watermark is just meant to tell you whether the text was generated by a watermarked model, not whether the ideas came from AI or something like that.
In practice, it is theater. Are they going to do this with the code output too? This is nonsense security theater for the low thinkers to have a sense that someone is in charge. When we all know nobody is in charge, anywhere.
It knows when several tokens are about equal vs. times where one token is vastly preferred. In the latter case, that’s usually code or math or something similar and so it won’t alter those tokens.
I wonder if that is entirely true. It is also in the AI companies' own best interest to be able to detect AI generated text so that they can avoid training on it ("Habsburg AI").
In the case of Anthropic it would also be entirely unsurprising if they've been lobbying the government to force everyone to do something in their (Anthropic's) own best interest.
> It is also in the AI companies' own best interest to be able to detect AI generated text so that they can avoid training on it
It’s also in their interest to demonstrate that they can be trusted and to show that they at least pay lip service to limit the obvious downsides of the tools they are selling. The use cases they sell to mainstream audiences are not affected by detection tools. The point of having a LLM do the work for you is that the work is done, and reliably. It does not matter if it is done by a LLM, and most of the time it is obvious anyway.
The EU mandate for AI watermarking is likely the first step in the direction of prohibiting AI for specific use cases. The pretext for outlawing (or at the very least controlling) the use of AI when the time comes will be something along the lines of data integrity or just general compliance legalese.
So, the use cases that are being sold to mainstream audiences actually will be affected by detection tools, especially if the output is intended to be monetized in some way. In the near future the EU will likely come down with heavy intervention to prevent AI from impacting employment rates across Europe. The number of legitimate use cases for costly frontier models drops significantly once eliminating professional jobs is off the table. This is all conjecture at this point though.
I’m sorry, this is going to be a bit long but you made good points.
> The EU mandate for AI watermarking is likely the first step in the direction of prohibiting AI for specific use cases.
I am not sure how practical that would be. The cat’s already out of the bag and they won’t prevent companies in the whole world from releasing open weight models. Playing catch up by distilling flagship models is also relatively cheap; we’d see smaller companies setting up shop in friendly regimes. And I don’t see any appetite to go full child porn and criminalise the possession of a LLM. So we’d end up with a similar situation as with illegal downloads, i.e., everyone will do it and nobody will care.
> The pretext for outlawing (or at the very least controlling) the use of AI when the time comes will be something along the lines of data integrity or just general compliance legalese.
They could forbid using LLM for hacking, but hacking is already illegal. They could make it a factor when determining punishment, but I don’t think that would work terribly well. Most of the dangerous stuff we can do with LLMs is already illegal, or should become so. Things like propaganda, identity theft, harassment, scams. We need enforcement with teeth on these, not pointless feel-good legislation. Again, there are parallels with cryptocurrencies and torrenting software. These things have illegal uses, but it’s also really difficult to make them illegal, at least in semi-functioning democracies.
> So, the use cases that are being sold to mainstream audiences actually will be affected by detection tools, especially if the output is intended to be monetized in some way.
I don’t know that mainstream audiences are really against LLMs. They are mostly against AI in a nebulous sense, but even non-technical people use ChatGPT or equivalent. I think that the critical mass is already there and the tools are convenient enough that they couldn’t outlaw them without a massive uproar.
Detection tools don’t seem all that relevant to mainstream audiences’ use of LLMs. AI companies will sell this as a safeguard against misuse, and everyone will be happy about it. The politicians will say they accomplished something, the AI companies will slowly turn public opinion, and the public will have shiny toys.
> The number of legitimate use cases for costly frontier models drops significantly once eliminating professional jobs is off the table.
I don’t know. They can open possibilities that we don’t necessarily consider.
One example I have is a friend who is getting his house refurbished. He’s not an engineer or a material scientist. He does not have enough free time to read thoroughly on the many subjects involved. With a decent LLM, he could untangle the technical documents sent by the architect and the contractors to really understand what was going on and be involved, rather than passively follow the architect’s advice. For starters, the LLM was very useful in finding issues in the quotes he received when he was looking for an architect. Those were long, technical documents, with no really standardised structure and full of jargon. I don’t think that person is going to want to stop using LLMs now. Many people are having this sort of moments right now.
> In the near future the EU will likely come down with heavy intervention to prevent AI from impacting employment rates across Europe.
Maybe. But i don’t believe the EU is well equipped for that. Labour laws are largely local and different in each member state. The EU regulations are basically the common denominator, ore or less, and it is easy to see why: for regulations to get adopted, they need a strong enough majority in the Commission, in the Parliament, and in the Council. It is very difficult to get anything controversial that affect the sovereignty of member states passed.
The angle of the current AI regulations is that they set the rules for the single market, which is where the EU is the most legitimate. It is difficult to see market angle for the effect of AI in labour, and I think enough member states would be keen to kill the project.
Also, there are many influences at play, but the EU is fundamentally an economically liberal institution. It very rarely goes in the direction that reduces economic activity. Look at how clumsy it is at fighting against cheap Chinese imports. I don’t think the institutions themselves would really want to make AI illegal. Companies have too much to lose.
Great points all around, lots to think about. It's totally possible that I am way off in my assessment of the future.
I think your comment about torrenting touches on an interestingly relevant case study, specifically media piracy. The seed of that technology was sown when the internet was still just clusters of machines passing files around and then exploded with the PC, Napster, and TPB. Napster tried to be a legitimate commercial enterprise with a business model of undermining the ability of copyright holders to rent-seek on the consumption of the material they 'owned'. At the core it was a novel technology (P2P) that revealed an economic arrangement to be out of date (if a distributor no longer has to manufacture a copy of the media for each individual consumer, their business model boils down to rent-seeking). Western legal systems were quick to rule on the matter (in favor of copyright holders), and I have no doubt that they were looking ahead to a future wherein media creation was totally disincentivized by said novel technology.
Now a novel technology (the cloud inference-backed LLM) is challenging another economic arrangement. This time around, the arrangement being challenged is the higher education->professional job pipeline. All advertising and media messaging aside, it really does seem like the frontier labs are only economically viable if they get massive enterprise deals across a broad spectrum of industry. There is fundamentally one chunk of capital organizations are going to spend either supporting their talent pipeline or padding it (to put it gently) with enterprise LLM deals. If the latter path is taken too far, consumer spending plummets (due to lack of middle-class incomes), assets backed by consumer debt/spending fail longterm, and we will have to deal with a deluge of socio-political issues stemming from the absence of real social mobility (we are in the early stages of this now, incidentally).
All this to say, I think parallel situations from the past can guide our thinking re: AI regulation and the forces shaping it. Reasoning based on regular political/economic incentive structures (e.g. economic liberalism not wanting to reduce economic activity) will fly out the window at lightspeed once fear becomes a factor. I personally think its great that LLMs empower individuals such as your friend to increase the control they have over real issues in their life; that is what technology should be doing for us. Cloud inference-backed LLMs are doing the opposite: drastically reducing the power that the everyman has over his socio-economic future by throwing high-paying career paths for a whirl and incentivizing powerful organizations to destabilize the labor market. Only time will tell how this plays out.
What happens if we train models (GPT or human students) using the outputs of a model with text havingbthose watermarks? Is there something preventing the watermark from being learnable?
In addition to holding the key, wouldn't you additionally need to know exactly which model to check against? So for passive detection to happen, I think each company would need to check every message against every model version? Also would need to spend resources re-invoking the each model version against each message.
Probably they bias the RNG for selecting the next token. This can be done practically in a lot of ways, including during training.
I suspect the signal will be significantly under the noise floor, so it's not detectable if you don't know exactly what to look for, but certainly you can submit more information then the textual contents.
But if the user's prompt is in the context, you don't know exactly what the RNG chooses between. I don't know what trick they use to get past that, but it seems impossible to get by it in the general case (i.e. if the prompt can be anything) and you'll probably quickly compromise quality if you try.
A reasonable guess about the algorithm is 'A Watermark for Large Language Models' (https://arxiv.org/abs/2301.10226). The idea is that each generated token (or bigram) seeds a strong PRNG that splits the vocabulary into a 'green' and 'red' set. The sampler then tries to select a 'green' next-token for generation.
After-the-fact checking only needs the vocabulary splitter, which is independent of the LLM. Over a sufficiently large text non-watermarked text would expect to use green and red tokens with the baseline probability, and that difference can easily become statistically significant over sufficiently long texts.
The basic algorithm has obvious knobs to tune, among them the initial ratio of red to green tokens and how hard the sampler tries to pick a green token. These would balance fidelity to the original distribution against watermark detectability (minimum required content length for statistical power).
Anthropic actually tells you the approach they use, and it's not that. From their Claude Text Watermark page[0]: "Claude’s text watermark is a version of the SynthID-Text approach published by Google DeepMind in a Nature paper in 2024."
The Nature paper is "Scalable watermarking for identifying large language model outputs"[1]. This method does not separate out tokens into separate classes, but merely uses a seed for the PRNG that selects which among the most likely tokens generated by the LLM will actually be output. This has the advantage that there's no green and red token sets, so no token is systematically favored or disfavored. If a particular token is overwhelmingly predicted to be the most likely candidate, it will almost certainly be selected, so the watermark doesn't affect that. Even if there are several choices of output token at a point that have similar probability of selection, the watermark doesn't systematically bias in favor of one token or the other.
This is actually a quite elegant method of watermarking that, contrary to people's fears, won't adversely affect the model output. The main concern I have with it is that it appears that you can't actually test the watermark locally, without uploading it to Anthropic. I'm not sure why that's the case, since there's no particular reason the watermarking key has to be private, except if you want to prevent others from generating text with their own LLMs that is watermarked to look like it's generated by Anthropic - but everybody wants their text to not have the watermark.
Simple version: In instances wherein the otherwise statistically chosen next word is a "toss-up", watermarking removes the randomness by imposing specific choices, determined by a key. This then becomes a detectable pattern when scanned with the key (stastically—detection itself is probabilistic).
>use it to store arbitrary information
No additional data is embedded. The range of available data is constrained by the text being generated (i.e. the sets of "next words" per text).
From what I’ve read, they won’t be imposing specific choices, but using a different (biased) RNG for those “toss-up” choices. With enough sampling, you could detect if the RNG was biased or not.
This is what I meant by "imposing specific choices, determined by a key". Maybe "impose" or "specific" were too strong in my attempt to simplify?
I attempted to clarify that the impositions themselves are not deterministic, by indicating that the entire process is still probabilistic.
Maybe Anthropic's explanation is simple enough [0]:
>When watermarking is used, choices are still made at random, but the source of the randomness is different. Instead of using an arbitrary random number generator to pick the next word, watermarking uses the key and a few words that come before to settle what word the model should pick. That is, the words that Claude picks are still random, but now, one can check the sequence of words and see if it’s consistent with the choices Claude would make if it was using the key. If it is, one can assign a probability that the text was generated by Claude.
If it stores unique information, by definition it can store arbitrary information because it can point to arbitrary information. So they can have it relate to anything they want. Even a full breakdown of the original text if they choose.
Removing the watermark usefulness depends on your use case.
If you care to avoid detection, yes, it is useful. If you care about the best possible sequence of words, then the damage is already done once watermarked.
Maybe those engineers that don't care about making their solution _good_ are actually the good engineers, have you ever thought about that??? There are exceptions where quality really matters but most software shops build CRUD web apps where it's not really a good characteristic to obsess over the cleanest and most elegant details, instead you need to get the stuff the customer wants done!
I for one tend to care less about the minutiae of solutions implemented by AI as long as it gets the job done, I do care about architecture and design decisions and correctness and I have ways to steer and verify these when working with LLMs but I couldn't care less about it writing "good" code. Bad engineers also produce better results with AI at least when they're working in established frameworks, AI doesn't really need a lot of high level architecture input when designing or building a web app with a common stack, so as long as you're not working on something that's completely novel I don't think it will make a strong difference.
Maybe designers think the same way about the AI generated web designs I have Claude Code do for me but to be honest I don't care, I just know that before this tool existed it would have taken me weeks or months to come up with a good design and I would have to rely on prefabricated UI libraries and stuff like that or pay a designer tens of thousands of USD to make one for me, now I can get a (for me and my customers) perfectly acceptable and professional design within a few hours. So maybe I'm also a bad designer that amplifies my bad design taste 10x in my company, but the fact is the stuff ships and makes money and the customer is happy! And I can tell you customers or users don't give a shit about how good your code is, they only care if the software works and does what they want!
Totally agree with your framing here, but that is also what I fundamentally consider "good" code -- it's code that solves the problem that your customers/business needs without making it _harder_ to solve the next problem.
There are valid situations where the best code you can write is code you never look at and throw out the next month; there are equally valid situations where the best code is well thought through and reasoned abstractions for an area you expect to become core to the business in the near future.
Thanks for the wording on your first sentence there. I've been trying to figure out a way to get that thought expressed succinctly.
I think with AI coding, what we call "good" code changes. Lots of abstractions really only exist to help load the context into the human brain so that they can solve the next problem. If an agent can just search and find all the places to make a change, or to duplicate code with small changes for the next problem, is that bad? Does is just feel bad because that's not what we're used to?
We use structured looping instead of gotos because that makes sense to us, but the compiler still turns it into jumps in assembly. If our interaction is now at a higher layer, do we need good "code" or do we just need good "architecture"?
I don't know, but it's just something I've been thinking about lately.
Using context always means there's less for something else, whether you're a human or a machine. Abstractions that localize reasoning and help load the context into a human brain are ideal for machines and humans alike.
In my experience, those that can’t understand design in the small (code level) don’t understand it in the large (sw or systems architecture) either.
Doing the right thing for the customer is independent from good design and good code. It’s a problem of requirements and project management.
This is an excuse some poor programmers use, that they can’t write good code, but at least they fulfilled the customer requirements :D
This is a great point. I find myself in the middle, where I appreciate and look up to past coworkers who were way better than me at attention-to-detail, but I have a tendency to focus on delivering actual results quickly with maintainable and easy to read code. A lot of engineering teams go into their own world on adhering to best practices with no eye on inefficiencies in the process that causes days of delay because it makes them feel better. I definitely believe in shift-left (findings bugs early saves a lot more time than finding them in production). I believe in balancing fast delivery with quality/process.
I tend to find that this perspective comes from people working on relatively small or isolated projects.
On a large system, the customer being happy today isn’t enough. You need other engineers to be able to understand the system.
Have you ever been on call and been woken up in the middle of the night to fix a production incident in a system you didn’t write?
If everything you build is small, isolated and easy to replace (basically fire-and-forget), then yeah... who cares? Ship the ugly thing, get paid and move on.
If you’re going to be working on something for the next 5+ years, you should definitely spend some time thinking about what you’re doing.
So I'm greatly excited what AI will bring about in physics, more so than in math, because in physics it's clear that our fundamental theories are missing a big piece of the picture, and given how easily AI crunches through Millenium prize problems I think it's possible that AI will come up with a viable grand unified theory uniting quantum mechanics and gravitation, or produce new predictions in other areas. There's enough contradictory or unexplained observational data available to make a ton of progress on the theory side I think. Exciting times ahead!
reply