Hacker Newsnew | past | comments | ask | show | jobs | submit | app13's commentslogin

Humans do that by inventing algorithms like AlphaGo :)

In a more serious note, I think the person you are responding to meant "new" more in line with "novel".

Take a microfluidic chip, for example. Current AI systems can create new flow cell geometry, but cannot come up with the idea for a microfluidic flow cell itself.


Last model from Anthropic that I can use as a security researcher. 4.8+ won't even look at my git repos

Granted they contain robot dog malware, but still.


The moment it sees shellcode, it runs for the hills.


Whats your setup? I have a single 3090 and am struggling to get it purring


5900x, 3090 24gb (slightly undervolted), 128gb ddr4, running via Ollama.

I am benchmarking it now locally, will put the results and speed/tps on aibenchy.com


How did you undervolt the 3090?


Msi afterburner, you go around 900mv curve editor, raise it up to normal clock frequency, and it uses like 260w instead 300w for same performance.

There's some YouTube guides for it.

I also undervolted my new 5070ti, same tdp, around 260w instead of 300w and like 8% better performance.


I remember the days of crypto mining memory intensive tasks do not suffer from TDP reduction as much as game rendering.It's so counterintuitive.I remember getting 96% performance at less than half of TDP with my old AMD card.


I needed to thoroughly test rerankers on my companies rather unique corpus.

Opus and I wrote a parallelized test harness and labeled groundtruth in around 2 hours.

In 2022 that would've likely been all I did for a couple sprints


I encounter this regularly and it still feels weird.

That sense that you did something better in a few days than you would have in a month 5 years ago. It's like buying a table saw for wood working.

One crazy thing I think about often is how there are so many correctness and testing harnesses that would have taken weeks to build in the past so we simply never would have. We'd just do our best then wait and see what comes to the surface. This is a huge part of what makes it possible to actually make better software with LLMs in my opinion. It isn't just 'LLM codes better than I ever could' (that's often untrue still) but 'LLM enables me to make assertions about the program to degrees that would have been absurdly impractical in the past'. It's huge


Yes 100%. This morning I casually prompted Codex to drive the browser to complete extensive performance testing in-situ that would have literally been weeks of work before. Probably in reality it just wouldn't have been done, and performance guarantees would have been attempted up front via more careful design.

In this case the design was also AI generated, and there were limited wins to be found because the design was already superb.


The CVE process itself is broken. HackerOne and company VDPs are inundated with new reports of varying quality thanks to the advancements (I think) in agentic AI. It's allowed for both an increase in trash-tier low quality AND legitimately high quality reports. Since the same AI's are writing both, its almost impossible to distinguish between the two at a surface level.

In response, companies just aren't responding like they used to. I spoke at a cybersecurity conference In June and the overwhelming "vibe" on the floor and in the talks was that responsible disclosure was dead or dying, and public disclosure is the way forward. The Microsoft and Nightmare Eclipse situation was oft cited.


As someone on the company response side of the HackerOne brokenness, I can confirm that this effect is real but would also note that the difficulty of distinguishing is not as severe as all that, because the companies have access to the source code which the researchers do not typically have access to.

This means that the token cost of verifying any given HackerOne report is dramatically lower than the token cost of producing a report in the first place. Automated triage systems should be possible, and realistically it's well within the capabilities of most companies to go further and actually automate the Red Team side of it and catch issues before they surface in the black box research. From what I've seen doing so should cost dramatically less in tokens than the bounty payouts do.

The problem is that security is woefully underfunded in most companies, so even an infosec organization that saw the deluge approaching from a distance may well not have had the resources to prep for it even if they knew exactly what actions they would take if they had the capacity.


The token cost of a report is lower bounded by the number of tokens in the report * price per token of the cheapest model. The token cost of a good report is much higher, but sifting out the good reports is the entire problem.


In theory, yes, but are you actually seeing clearly-garbage lower-bound reports like this?

The ones we're seeing show clear evidence of being AI-generated, are often incorrect or duplicated, but they also show clear evidence of the AI having done its homework and spent a while crawling our API.

Even if we were getting reports at the lower bound you're describing, those would be even easier to triage: just add a quick step to check if the API in question even exists, then if it does that very cheap "where is this API" query becomes part of the input to the second-level triage that spends more tokens.


Do you have a tier based system for identifying high value reporters? Like a credit score for vulnerability hunters?


Not formally, but we do informally recognize names and prioritize researchers who reliably turn up good results.


that might be a nice incentive for bounty hunters, a sign of recognition and a good metric for triage


Their CISO literally acknowledged it and then they all continued ignoring it again. This isn't just bad process, this is a broken security organization.


Should a company promoting the enterprise usability of AI, itself start with building a intake process to distinguish between the noise and signal for these reports. If you cant solve your own problems with your product then how do you expect the customers to be able to use it.


Not even that. Even before AI came along the widespread practice of CV-Enhancement was slowly strangling the reporting of actual legitimate, needs-to-be-fixed issues. When it turns into a giant shit-shovelling exercise it's not surprising that some of the shit doesn't get shovelled.

Not defending HackerOne, but pointing out that it's not a black-and-white issue.


I will be totally honest, it is hard for me to want to engage with content that is clearly written by AI.


I don't know how people can even manage to read through this stuff, much less upvote it. It's even more depressing to me that a comment her got flagged for pointing out this was slop.


I participated in research from 2017-2022 that found similar results regarding bio-interactions, generally.

Learned a lot about making microfludic flow cells at least


Bravo to the Domo CDO, who's fomo won't stop slo-mo no-mo'


Yolo Domo Arigato


Mr. Roboto


Oh...


Another Chuck fan I see. I'm glad to see I'm not alone.


Domo CDO says no mo slop yo


won't stop with no slop slo-mo no-mo'


Routing in agent pipelines is another use. "Does user prompt A make sense with document type A?" If yes, continue, if no, escalate. That sort of thing


For this type of repetitive application I think it's common to "fine-tune" a model trained on your business problem to reach higher quality/reliability metrics. That might not be possible with this chip.


They say LoRA finetunes work.


$25-35/hr for the evening shift it seems. Not starvation wages by any means, but certainly not great (or even good) pay for that area.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: