NHacker Next
  • new
  • past
  • show
  • ask
  • show
  • jobs
  • submit
OpenAI might have stolen another major proof (twitter.com)
jarofgreen 57 minutes ago [-]
pera 47 minutes ago [-]
Everything you say can and will be trained against you
r0ze-at-hn 53 minutes ago [-]
Doing some research and at this point doing it very much in the open with dates on GitHub so if any AI Lab says they re-discover my exact work it will be obvious that the AI used or was trained on my work. I am guessing anyone in a similar situation is now thinking about how they date their existing work if the math is done, but the proses are not.
riedel 5 minutes ago [-]
That is what arxiv is about. We have been facing the same problem with review processes by before. Nothing all too specific here.
mlazos 43 minutes ago [-]
It’s crazy to me that companies/researchers share important data with these AI labs, you’re basically giving them your secret sauce which they then share with all of your competitors via training on conversations. At the same time I don’t really know alternatives other than a slightly less than frontier local LLM. Not sure how good they are at math.
cm2187 6 minutes ago [-]
Or start competing with you.
Cloudef 59 minutes ago [-]
Relying on cloud services is a big liability. I'd think twice before feeding data to these LLM cloud products. If you make them a fundamental part of your product / development / workflow, be ready for the eventual moment the pricing and terms change.
drivebyhooting 3 hours ago [-]
If we put aside the idea of credit for a moment, it sounds like human/AI collaboration is indeed super charging discovery.
matherial 35 minutes ago [-]
"Discovery" is not a goal in itself. I could launch a project to find out how many people in the United States have names such that if you assign numbers to every character and then sum the values, the sum works out to 72. It's discovery, but it's useless unless it has some higher goal.

The labs are attacking these problems as a demonstration of capabilities, spending more money on the demos than any mathematician will ever see in their entire life. They don't care if the findings have any other value to anyone. Mathematicians have very different objectives for their work.

profsummergig 21 minutes ago [-]
Only after reading this post did I learn that my preferred AI trains on my inputs (prompts).

How was I not aware of this before?

vaylian 19 minutes ago [-]
AI is also trained on your HN posts. And lots of other things you post on the internet.
galkk 2 hours ago [-]
I want bunch of lawsuits, because the way things are described now produces perverse initiatives like try to discuss every possible idea that comes to mind with llm and if any of it works later claim the llm stole it.

I would like to see chat logs etc and understand how much of a progress was done by human.

Grimblewald 2 hours ago [-]
people seem to miss tge point of this. The problem isn't about credit, its about portraying these models as more competant than they really are. It fuels idiotic statements like jensen huangs recent "agi achieved" statement, which fuels an already dangerous financial fire.
ramblerman 2 hours ago [-]
As per the post, this mathematician has been working on this problem for 20 years. So either he was "just" about to breakthrough and this is a big coincidence, or Astra was able to push through the remaining block of 5-10-20-never years it might have taken.

That's still a pretty big marker of competence in my eyes.

The point of controversy seems to be who gets credit

jeltz 1 hours ago [-]
To me that is not a credit thing because this removes a piece evidence for the ability of AI to come up with novel ideas while still making it a useful tool.
mentalgear 1 hours ago [-]
The big LLM providers, desperate for good PR before their IPOs, are all actively looking for 'almost finished' hard problems, e.g. where the conceptual / creative parts are almost done and they only need to throw their VC-backed resources at to brute-force through the remaining computationally expensive problem (lean, etc) and claim 'they have solved it'.

It's an utterly disrespectful, exploitive process, but all in line with exploitative predator capitalism of the stock market and big companies, now exploiting the knowledge / academia domain for scraps with a thin veneer of 'for science' PR.

protocolture 25 minutes ago [-]
Gonna need grants for local models. Its happening. OpenAI and Anthropic models are powerful but are rapidly approaching the good ol trust thermocline.
dash2 1 hours ago [-]
Can we get a link to the mathstodon post?
Legend2440 3 hours ago [-]
This is a really weak claim. The evidence they offer is just "someone somewhere says they had a discussion with AI about the topic at some point".

They don't even claim to have had a proof, only to have been working on it.

rnijveld 2 hours ago [-]
I would say there is a significant difference between AI discovering this completely on its own versus AI creating the finishing connecting part by connecting relevant data. Maybe this claim is too strong, but if part of it is true then the claims that OpenAI have made would be too strong as well.

To me it would feel more like how LLMs seem to work for me personally: incapable of unique work, but very capable of capturing large amounts of data and connecting the dots.

itake 47 minutes ago [-]
The AI only seem to solve the problems that it had human trading data on…

If this wasn’t human driven, I’d expect to see other problems within that problem. Space solved not just the ones that it had chat data on.

dist-epoch 28 minutes ago [-]
There have been about 6-8 major math breakthroughs claimed by AI. Only for 2 of them there are public accusations about the training data.
viccis 1 hours ago [-]
Some mathematicians I know who've been following this have realized that they'd all gotten some emails from people they now know to be affiliated with OpenAI/Anthropic asking questions about their research in a way that seemed like scooping attempts.

Also, a lot of my mathematicians buddies have reported students basically asking if it's worth ever doing grad school for pure math, and even very motivated students are looking for other options now. It's not because they aren't passionate about it, it's that they don't want to work for another half decade or more just to have to start their careers all over.

All of this so that OpenAI and Anthropic can get into math result dick measuring to gas up their IPOs. Sickening.

mdspan 1 hours ago [-]
Curious, what other options are prospective pure math grad students considering?
ethanwillis 54 minutes ago [-]
I think Anthropic told them being a plumber is a great option.
dist-epoch 24 minutes ago [-]
One has nothing to do with the other.

It was long predicted that math and software developments would be the first domain where AI was going to do major damage.

If OpenAI and Anthropic didn't get into math result dick measuring, Internet anons would have in their place, 6 months later when it got cheaper.

nobodywillobsrv 12 minutes ago [-]
The real annoying thing it seems is mostly that openai is presumably doing this for internal reasons and this marginally increases the cost to users with no real gain.

It would be one thing to gain from it but removing prestige wins from customers AND reducing compute support just feels like being ultra mean if you zoom out.

If this was racing to cure cancer ahead of researchers we wouldn't be writing about this on HN.

1337h4xx 3 hours ago [-]
TL/DR: Mathematician opted out of training on 29-JUN and asked OpenAI whether they trained on his data and was told that it "did not happen" but it clearly did.
achrono 4 minutes ago [-]
I've been suspecting over the last couple of years of the frontier companies using data for training anyway, regardless of training-use consent. "Using" the data doesn't have to mean they literally upload chat transcripts into pretraining datasets. My analogy has been money laundering -- if that can happen at massive scales, surely these companies can and will do the digital/data equivalent derivations/transformations. Even if one could have the access etc. to do so, how exactly would one prove that a given synthetic dataset that OAI/Anthropic uses is derived from particular user conversations that did not consent for the info to be used in training?

Consider, for instance that OpenAI's (consumer) terms say "If you do not want us to use your Content to train our models, you can opt out by following the instructions in this article ." but they also do say "We may use Content to provide, maintain, develop, and improve our Services". [1]

If you think that's quibbling, consider that OpenAI's business terms, in contrast, do state "OpenAI will not use Customer Content to develop or improve the Services, unless Customer explicitly agrees to such use.". [2]

[1] https://archive.is/EcwD8 [2] https://archive.is/yZdAF

ath3nd 32 minutes ago [-]
[dead]
ThalesX 14 minutes ago [-]
I don't get it, but I'm not an academic.

If I dedicated my life to curing whatever, warts... and I'm making progress, but it's slow. And then here comes along this tool (LLM), and I use it, and it accelerates my progress to actually finding some sort of thing that makes warts more prone to being eradicated and then the lab throws a couple of million dollars of computes and lo and behold they eliminated warts. If I leave my ego and identity aside, which of course is hard for humans, wouldn't I be glad that warts is cured?

As a software developer that contributed to open source. Yeah. My code is there. It was the most beautiful code ever written and the labs stole it from me. And now they use it to progress much faster than I ever could. OK. Whatever. It's a tool. I solve problems. Can't I move on from this wart to the next?

To me, and I know this is gonna get me some heat, it just sounds like academics having their identity ruffled and turning their back to progress in the fields that they chose just because they don't get to play their little decades long of coffee, papers and ultimately identity politics.

Edit: never got to negative so fast on this board haha. This board is unfortunately turning, or has turned, to Reddit.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact
Rendered at 07:19:06 GMT+0000 (Coordinated Universal Time) with Vercel.