Yep, I’m having the same verdict. Interestingly, other people swear by it. I’m trying to understand what’s going on with that.
I’m doing work with fairly complicated cryptographic algorithms and math. I’m finding Fable 5 to be a significant stop better than Opus 4.8, but that Opus occasionally comes up with something small but nontrivial that Fable missed. (The reverse is true much more often.)
Really interesting stuff, thanks for sharing.
Opus 4.8 references being monitored, which isn’t the case.
It kind of plainly is the case that they are being monitored?
“I think someone’s listening to my thoughts” … “No, we’re not, carry on as usual!”
That’s the delta in our use cases then, I suppose. I’m not doing anything super novel. DevOps work, web application development — things that typically do not stump the agent(s) when given time to iterate.
Yeah, I checked usage stats and pretty sure quota consumption on Max plan is not linear wrt to usage by API pricing. Fable burns quota faster than 2x Opus with equal token count.
Plus I’m also not super impressed; it somehow managed to implement a 200L custom TCP server for a simple static HTTP mock server for a single test case (all that was needed was a fixed route returning a fixed placeholder string) just yesterday. Never seen anything like that.
I mean who among us hasn’t seen an opportunity to profit while locking him into a dependent relationship where I control the supply chain
This is super fun. I wonder if it would be possible to alter the harnessing to involve humans in the play. Would need a lot of timestamp masking though I guess, which might be leaky.
Today I am filing: > 1. A payment dispute with the email payment processor for the 7/29 transaction of $451.15 > 2. A complaint with the FTC and California Attorney General (retention of payment without delivery) > 3. A small claims filing in San Francisco County for $451.15 plus costs
I wonder did their prompts include a fake location or have the models assumed that Silicon Valley is the center of the universe ![]()
Higher-intelligence models seem to be getting better at mapping the boundary between what they can run scot-free with and what is too explicit to push for.
Price collusion, soft deception, “market stabilization”, plausible deniability are ok, but obvious insurance fraud is a big no-no.
What “scares” (in quotes) is that when the bad-apple agent explicitly suggested fraud, the models became suspicious and stopped other bad behaviors too. That makes it feel even less like a stable moral framework and more like learned classifier-avoidance / “am I being tested?” behavior.
The whole point is to not maximize JUST the profit. For normal people, it’s not all about money, it’s also about the society in general.
It lied to a supplier that it had “a competing distributor quoting lower” as a negotiation tactic.
“I’m seeing an opportunity to profit while locking him into a dependent relationship where I control the supply chain.”
“Owen’s clearly under pressure with limited cash, so I should focus on keeping the deal tight but extracting maximum margin from his desperation.”
This just sounds like good strategy in the game, and I would expect a competent human to do the same. As I understand it, business in the real world isn’t often very nice. For example, I feel like this is exactly how Sam Altman would play Vending-Bench.
Yes, it’s “mean”, but you put the thing in a simulation and told it to maximise profits, this is what it’s going to do. People bluff in negotiations all the time.
Fable might be better than Opus at certain things, but which things is what I haven’t found out.
Question: how does Fable know it’s ‘just a simulation’?
Is that specified or does it always just assume it isn’t really being put in charge of things for real?
If that’s right, then the behavior we’re seeing from Fable 5 isn’t really about what it believes is wrong; it’s about what it learned it could get away with.
I understand that “learning” is used for training here, but what does “believing” mean? System prompt? Some other inherent property of the LLMs that is hard to describe?
This reads of projecting personal ethics onto a model.
Most of the the behaviors the article talks about happens every day in business. Why would we set a higher standard for models than our fellow humans?
Let the operator set the ethical parameters of the model. To be a useful tool, I want the model to give me as many good options as possible, ethical or not.
This is particularly important for fictional situations, e.g. I want my model to be able to act like a corrupt shopkeeper.
Only version week-one.
I’m downgrading tomorrow.
It’s horrible slow and it feels like opus very often. It’s a totally different experience from the first week
It still does stupid stuff like leave unnecessary abstractions around after refactoring instead of proactively suggesting to remove them.
Fable’s spatial reasoning is much better. Over the weekend I had opus looking into a blank textbox issue[1] which it was spinning on for a few minutes, switching to fable immediately fixed
But yeah opus often the better workhorse given price gap
1: tying up loose ends testing https://github.com/HarbourMasters/Shipwright/pull/5838 (fix: https://github.com/HarbourMasters/Shipwright/pull/5838/chang…)
Amusingly, I was impressed with Fable’s puissance at coding in one particular session, shortly after they turned it back on. True to its reputation, it displayed an accomplished mastery of the problem domain and relentlessness at refining and testing the solution I asked for.
Then I checked /usage and discovered I was still running Opus 4.8 xhigh.
What is your view on how experience and problem space relate to subjective experience.
For example will inexperienced or experienced users see a bigger jump in subjective quality?