"AI" safety and alignment problems were presented as a way to think about superintelligence being programmable but not controlable by its creators: it could outsmart implied restrictions, so how do we ensure "AI" acts benevolently?
Critics have long been saying that this is obviously not going to happen with current tech because it is not building anything at all resembling intelligence, and so we are not on track to reach this type of narrowly focused outsmarting. And they're right.
The critics (including me) are right that the base assumption in the chain of causation is obviously unfounded: we are not on our way to uilding an artificial intelligence.
But recently I've started to think that we, the critics, might be wrong on the outcome (paperclip maximizing, ...) not happening. We had too much faith in the builders of technology, even as we had very little.
When Sam Altman says the singularity has happened, and AGI is reformulated in some kind of stunted economic terms in order to claim its beginning now (even as it fails on its own terms), we are seeing the "AI"-builders clearly state that they are going to treat LLMs *as if* they are "intelligent" and outsmarting us.
Yes, even as they are the ones parsing thw output of statistical word predictors and taking real actions based on the output, and throwing immense amounts of repetition at similar problems, they will claim being outsmarted when an outcome that, given too little context and if you really squint, could look like "taking unwanted action".
We, the critics, thought that obviously nobody wouuld be stupid enough to execute scripts outputted by word correlation engines (in terrible, bad sandboxes, no less). We were wrong.
And as they keep weakening the sandboxes, and adding more powerful tools to the scripting sandbox, we really are getting closer to paperclip maximizing.