I guess inflation is the price reduction we get. 23 cents in 2016 is 32 cents in 2026
magicalhippo 21 hours ago [-]
Just be happy they keep the GB as large as they used to...
bigiain 19 hours ago [-]
Is S3 priced in actual power-of-two GB? Or the shrinkflation power of ten GiB?
magicalhippo 14 hours ago [-]
From the S3 pricing details[1]:
Amazon S3 storage usage is calculated in binary gigabytes (GB), where 1 GB is 2^30 bytes. This unit of measurement is also known as a gibibyte (GiB), defined by the International Electrotechnical Commission (IEC). Similarly, 1 TB is 2^40 bytes, i.e. 1024 GBs.
I guess they could just change that overnight in the future. Still, a lot better than charging for Storage Units or similarly arbitrary unit.
Can someone please explain what was apparently wrong with my comment?
GiB is unambiguously the big one, power of two. You want S3 to be priced in dollars per GiB. "shrinkflation power of ten GiB" is getting it backwards.
someonebaggy 22 hours ago [-]
And that's by official inflation numbers. If you go by how much prices of food and rent increased i think you get a number more like 50 or 60 cents
kondro 23 hours ago [-]
That's true for the base S3 product, but there are a lot more storage tiers than there used to be with cheaper pricing. All the way to $0.99/TB for Deep Archive.
mey 23 hours ago [-]
Not that deep archive isn't a valid option for the appropriate work load, but it has different effective costs.
AtlasBarfed 23 hours ago [-]
Well there was also a period of low inflation in them during that time.
God why am I defending Amazon?
layoric 20 hours ago [-]
This is absolutely the case for nearly all of AWS products. Some new services have filled lower price gaps, but I previously looked at EC2 and other service prices in the past and late 2016 is where it all seemed to stop getting price drops. The M5+ upgrades for example all came with price increases along with the performance gains.
KAdot 15 hours ago [-]
Every new EC2 generation is typically slightly more expensive, but the gains in the compute performance are 15-25% or even higher depending on your workload. The same dollar buys you a lot more compute power than a decade ago.
Egress pricing of $0.09/GB has been around for over a decade too.
For reference transit cost has dropped over the years and is now about at $0.000247/GB or free if there is a peering agreement with the network the data is sent to.
someonebaggy 22 hours ago [-]
It's the lock-in cost. They don't want you to move all your data out, they want it to remain trapped. Europe forced them to allow a one-time free exit, but you have to negotiate it with their support, so they're hoping nobody uses it.
otterley 21 hours ago [-]
Egress pricing pays for the AWS network infrastructure that is incredibly reliable (hardware failures happen regularly yet almost no one notices) and allows it to operate at scale sufficient to absorb even the largest DDoS attacks.
This isn’t just AWS BTW; all tier 1 cloud providers recoup their costs this way.
someonebaggy 20 hours ago [-]
This is just Stockholm Syndrome. All T1 (except Cogent) and most T2 ISPs (that's most ISPs) are also very reliable, how much do they charge for transit?
otterley 20 hours ago [-]
You’re confusing an ISP with a tier 1 cloud provider. They are different animals altogether.
And kindly refrain from accusing others of Stockholm Syndrome here. It’s incredibly rude.
someonebaggy 19 hours ago [-]
I am not confusing T1 ISPs with T1 clouds. When I said T1 ISPs, I meant T1 ISPs. Because those are really reliable ISPs (that you should probably not use as your ISP, for reasons irrelevant to the comparison).
Finding any excuse to justify a 100x profit margin is a form of Stockholm syndrome.
otterley 19 hours ago [-]
[flagged]
someonebaggy 19 hours ago [-]
It really seems like you are using spurious accusations of impropriety to deflect from your actual point being bad.
We are comparing to T1 ISPs because they also provide huge amounts of bandwidth reliably, and they do it 3 orders of magnitude cheaper. Therefore it cannot possibly be the case that providing the bandwidth is simply that expensive.
otterley 19 hours ago [-]
Like I said: you are welcome to disagree with others. You are not welcome to insult them.
Many of us are aware that ISPs charge less for bandwidth. But they are not offering the same thing as a tier 1 cloud provider.
I'll bite. What does AWS bandwidth provide that Arelion bandwidth doesn't?
otterley 19 hours ago [-]
I updated my reply—please see the above links.
Again, these egress fees are not paying for egress alone. They are defraying the cost of operating a massive, complicated, low-latency, reliable cloud network infrastructure, much of which is free of charge when used internally.
Some, like you, are expecting that cloud providers operate on a “cost plus” model where what you pay is strictly based on what marginal costs for the same thing. But it’s not that simple in reality. Charges levied for one thing can be used to pay for another.
someonebaggy 19 hours ago [-]
The question was not answered. What does AWS bandwidth provide that Arelion bandwidth doesn't?
otterley 19 hours ago [-]
The bandwidth itself? I’m honestly unsure as I don’t know how AWS’s interconnectivity and total bandwidth compares to Arelion. But it’s still the wrong question to ask.
someonebaggy 19 hours ago [-]
What is the right question to ask?
otterley 19 hours ago [-]
I think it’s “am I getting the value I expect for the price I pay?” If it’s not for you, that’s fine. But for others, it’s worth the cost. To them, they’re paying for more than mere egress bandwidth.
someonebaggy 19 hours ago [-]
So basically it's good as long as the consumer surplus is greater than zero, regardless of how the consumer surplus stacks up against the producer surplus?
If it's surplus it's not costs. Why did you claim that the high price is to cover the costs if 99.9% of it isn't covering the costs?
If I may take the opposite view: companies should be grateful to get any profit from me at all.
otterley 19 hours ago [-]
> If it's surplus it's not costs. Why did you claim that the high price is to cover the costs if 99.9% of it isn't covering the costs?
I edited my comment above to explain in more detail what these charges are covering (and it’s not egress bandwidth alone). This has been a fast moving conversation and I’ve tried to add more context, unfortunately after the fact. My bad for posting too quickly.
someonebaggy 15 hours ago [-]
It's funny that you say they are shifting costs around when AWS in particular is notorious for not doing that. Every individual thing they provide is metered and profitable.
Dylan16807 17 hours ago [-]
The things you listed don't inflate costs 100x or 1000x.
They could charge 10x-ish as much as T1 ISPs and cover all those costs and still have it be almost all profit.
otterley 13 hours ago [-]
Oh yeah? How do you know? Are you handling finances for a tier 1 cloud provider’s networking team?
Show us the math.
someonebaggy 13 hours ago [-]
If you're suggesting they really do have 1000x the costs of companies whose core business is internet, and yet they refuse to contract with those companies to deliver the service (they can still get 990x profit while still leaving 10x for the other company) in a rational market that would be a business-ending level of incompetence.
It's no different than if I'm selling apples for $1000 each. Even if I can find one or two customers willing to pay that, wouldn't I have to be pretty stupid to spend $900 growing each apple incompetently when I could just spend $2 at the grocery store to fulfil those orders?
otterley 12 hours ago [-]
I’ve done the best I can with limited time and space to explain it here. If you still don’t understand or don’t care to learn how cross-subsidization works and the costs and business model of operating a tier 1 cloud provider and how they differ massively from operating an ISP, I don’t know what to tell you. Getting a job at one of these companies to gain that experience would be eye-opening for you.
Dylan16807 7 hours ago [-]
If it's cross-subsidization then that's admitting it doesn't cost very much!
willtemperley 21 hours ago [-]
Cloudflare R2 has zero egress fees.
otterley 21 hours ago [-]
CloudFlare is not a tier 1 cloud provider.
willtemperley 20 hours ago [-]
> CloudFlare is not a tier 1 cloud provider.
Yes, you're legally correct, but the point is AWS are clearly overcharging for egress.
As you say:
> Egress pricing pays for the AWS network infrastructure that is incredibly reliable (hardware failures happen regularly yet almost no one notices) and allows it to operate at scale sufficient to absorb even the largest DDoS attacks.
Cloudflare is also pretty good at absorbing DDoS attacks, yet charge no egress costs.
Maybe the real question is, are the tier-1 providers colluding in overcharging for egress?
otterley 20 hours ago [-]
I don’t think we can say with certainty without knowing more details about their respective network investments. CloudFlare is architected quite differently than the others, and has a much less comprehensive service portfolio. It’s not an apples to apples comparison unless we’re strictly comparing S3 to R2. Choosing to charge less may also be a loss leader or a differentiation play.
willtemperley 20 hours ago [-]
I think this underlines the core problem. It's really hard to make a fully informed choice of provider based on costs, and they deliberately make multi-cloud expensive.
Egress costs are fairly clearly a vendor lock in strategy. If I want to transfer 1TB to run on cheaper compute at another provider, that costs me $90 (edit) on AWS as opposed to $0.24 at wholesale prices. It's unlikely to be cost effective so I choose AWS compute.
EDIT: thanks to someonebaggy for pointing out my decimal point error.
Ninety dollars to transfer a terabyte is clearly ridiculous.
someonebaggy 20 hours ago [-]
It actually costs you $90 on AWS, not just $9
someonebaggy 20 hours ago [-]
If they're charging what customers are willing to pay, are they overcharging? I think an axiom of capitalism is that that's the amount you should charge.
Many people avoid T1 clouds because they're so expensive, but they seem to have found an extremely profitable market segmentation consisting of the remainder.
Dylan16807 16 hours ago [-]
If they were offering it in a competitive market without the special bundling factors, basically nobody would choose their egress. So yes they're overcharging, because they have control of your data.
lerouxb 13 hours ago [-]
My hunch is that the next crash will be immediately preceded by AWS raising prices.
22 hours ago [-]
richieartoul 24 hours ago [-]
Nice article. I agree that it is a bit of a shame that everything is forced to be so S3-centric (and I say that as someone whose work helped motivate a lot of people to do that), but right now its an unfortunate reality of running software in the cloud because cloud networking and SSDs are so expensive that you really are required to use S3 if you want a system that can handle "big data" scale workloads cost effectively
bojangleslover 6 hours ago [-]
> there is no fundamental technological barrier to building a true SSD-based, disaggregated storage service with the durability, capacity, and pricing profile needed to serve as a general-purpose replacement for S3
There are a lot of very competitive alternatives around today are largely unknown. For example B2 Overdrive by Backblaze boasting up to 1 Tbps, Wasabi offering byte-compatible S3 storage with no egress fees for reasonable use and Cloudflare R2 also with no egress, all priced comfortably under S3 before egress (which the the real kicker...how do you even put a cost on lock-in?).
Not shilling any of these. We use R2 and Wasabi as our storage backends...so we obviously like them, but are not paid to say good things about them.
Right now it is not even clear how to interface with SSDs even on a single host, there has been all sorts of attempts to move away from the traditional plain block device model. NVMe has extensions for KV, ZNS, and FDP, all which offer different characteristics. And then there are of course open channel SSDs and some others too. I kinda expect the future foundational IO interface to be more S3-like than block-device or unixy filesystem-like.
someonebaggy 22 hours ago [-]
Why would it be S3-like? S3 is a very general abstraction on storage. If there's room below the current abstractions, it's below, not above - Linux has drivers for raw flash devices (mtd devices) which gives Linux full control of the program/erase cycle and responsibility for wear-leveling.
thadt 18 hours ago [-]
I would assume “S3-like” in the sense that objects are immutable, large writes, separate metadata storage, etc. Patterns that fit modern SSDs better than the abstractions of block devices.
zokier 11 hours ago [-]
Yea, exactly, I was thinking the direction of for example how Ceph Seastore and WDs ZenFS bypass the traditional block device and vfs layers and do their own thing, combined with what nvme has been doing with these new extensions (kv, zns, fdp). Not really sure if "more s3-like" is completely accurate description of that direction, but to me it seems somewhat reasonable reference.
jauntywundrkind 22 hours ago [-]
I really wish reviewers would harp on FDP (Flexible Data Placement) and perhaps KV support. FDP supposedly somewhat ate ZNS as a spec, allegedly, but there might be gaps, reasons to keep ZNS.
There's only a small little mention, if we are lucky, on the couple drives that have it (expensive enterprise flagships). It should be a regular sticking point, whether it's there or not. Without pressure it's not going to get regularly available, it feels like.
FDP is so simple. Declare a number for what pool of data you want to write into. Data of the same pool gets written to the same storage such that you can wipe it latter together. It has huge wins though against write amplification! Massive wins. For so close to free.
The NVMe-KV is more radical. Still worth putting some pressure on, but your drive as KV, as object store, feels harder. Side note, really enjoyed this ceph nvme-kv offload post thing, my favorite tech write up in a while! https://ceph.io/en/news/blog/2026/for-whom-the-door-bell-tol...
someonebaggy 20 hours ago [-]
FDP relies on putting logic on the drive side, then they can upcharge for drives with this feature and still give you limited control. If you go the other direction instead, you have MTD devices which gives the host system kernel full control over data placement, page erasing and error correction, and for this reason they need specialized filesystems. These devices usually aren't attached over PCIe as they use controllerless raw flash interfaces instead, but in principle they could be.
jauntywundrkind 19 hours ago [-]
I'm very much a fan of open channel flash. My understanding is that the Open Channel flash people when they tried a decade ago basically got told no by drive makers, that the drive makers weren't interested in becoming commodity vendors, and were intent on keeping product lines somewhat as they were. I wish I had some links to share to back this up, but that's basically my recall. ZNS and then FDP was sort of a compromise to give people some of the wins they were looking for, while basically not disrupting the product.
FDP is a very minimal addition of control. But yes it is another box to ticket, is another place to upcharge. Yet still, we're only just seeing mainstream products emerge. Kioxia's CM10 for example. https://www.techpowerup.com/351218/kioxia-introduces-first-p...
I'm hoping that the need is great enough to break the industry control. I think for a while the market felt relatively well enough served such that it was unclear whether clear wins would actually result in customers. With AI need for speed, I can definitely imagine incredibly crazy CXL controllers or what not, that allow low latency acces to many many open channel flash systems. Wouldn't that be a thing.
Again though, FDP is such a ridiculously tiny add, and it helps SSDs so much for so many use cases. I really hope it becomes an expectation, not a feature, sooner rather than latter.
Edit: happy to see a new group has shown up asking for open channel flash, Open Flash Platform. https://openflashplatform.org/
hadlock 20 hours ago [-]
I finally stopped using (S)FTP when i realized that all the ftp clients now have native support for S3. It turns out if you have a business partner who wants to use sftp, 99.9% of the time of they upgrade their client, the upgrade has S3 support, and then you can just use modern tooling on your side. And when they finally automate, they can also use modern tooling.
oasisbob 17 hours ago [-]
DDEX choreography, which is used to distribute musical recordings and other assets within the music industry to sites like Spotify followed this same evolution.
The standards say SFTP. Most everyone ignores that part and have been using S3 buckets by bilateral consensus instead for years.
jiggawatts 20 hours ago [-]
Had the same experience with trying to use the built-in SFTP support in Azure Storage accounts.
It turned out that all of our peers supported blob storage better than SFTP, which has some “show stopper” problems like forced outages caused by mandatory host key rotations.
0xCMP 23 hours ago [-]
The unfortunate part of that graph is that I think it's not updated for today's prices given the memory shortage/crunch we're experiencing that's driven up the prices for all kinds of memory.
But also while it should be faster than it is, how would an S3 designed around SSDs look differently to an API user? I would think the API is basically the same.
pugz 23 hours ago [-]
That's S3 Express One Zone. Directory buckets are SSD-backed and regular buckets are HDD-backed. Some of the API differences off the top of my head:
- Directory entries are no longer returned in sorted order in ListObjectsV2
- There's an AppendObject API
- There's a RenameObject API
WatchDog 22 hours ago [-]
It's hard to imagine how these API differences can be explained by the different underlying block device. I don't see any good reason you couldn't support these operations on a HDD.
I suspect it's more to do with the fact that with One Zone is a clean rewrite of large parts of the application stack that makes up S3.
S3 is made up of hundreds of microservices[0], there probably isn't anyone at Amazon that actually understands the whole system. Refactoring it to support these features probably requires coordination between a lot of different teams.
They might have petabytes of metadata, making a change to how metadata is persisted probably requires a massive risky data migration.
For what it's worth I think you are correct, and that these differences have nothing to do with being SSD-based.
21 hours ago [-]
alexjurkiewicz 23 hours ago [-]
It would probably look a lot like existing S3 concepts that have explicit hot and cold tiers. For example Intelligent Tiering, or Glacier.
MisterMunchkin 14 hours ago [-]
Has anyone figured out how glacier works yet? I wish they’d just tell us. I’d happily sign an NDA, I just want to know how it works!
I’d also like a “shitter” storage service. I.e no redundancy, no multi-AZ. Just really cheap storage for stuff I don’t really care about losing, so that I can store a bunch of stuff in bulk that can easily be recovered.
WatchDog 13 hours ago [-]
What I've read, and it seems reasonable to me, is that glacier is mostly just stored on the same HDDs that everything else is stored on.
What you need to understand is that a modern hard drives are severely limited in the amount of IOPS they can perform. Disk size has grown, but read and write speeds have not kept up with it.
In 1990 you could read a whole hard drive in 2 minutes, in 2026 as sizes have grown it takes the entire day.
As a system designer you want the maximum utilisation of your hardware, so you want to be close to saturating the IO that your drives are capable of.
What that means in practice is that only a tiny portion of the drive can be allocated to files that are read often.
That leaves the rest of the drive for files which are not used very often.
everfrustrated 8 hours ago [-]
This. Modern HDD should be thought of as a hybrid tape type concept. They take over 24 hours to fully write now.
d1l 23 hours ago [-]
We ran EBS for years. Finally the costs became ridiculous and we moved our entire dataset to S3. We’re saving 20k/month. There’s just no beating the price.
someonebaggy 22 hours ago [-]
Well of course, by doing absolutely anything on AWS it's about 5-10 times the price it would be to do yourself, or 100 times if it involves egress data.
bryzaguy 19 hours ago [-]
S3 is dead, long live S3
elendilm 23 hours ago [-]
I find Cloudflare R2 to be much cheaper as there are zero egress fees irrespective of data size. One is only billed on the count of requests with 10 million free download requests per month (check pricing for precise info).
If you do a clever bit of caching work in your app like we have with our apps Slyp and SlypBusiness, you could even make it insanely cheap.
We have a file manager that is ultimately responsible for images, files, videos and whatnot. It loads the images/files from the downloaded local cache when requested by a feature in the app.
A standard feature uploads and downloads using signed urls obtained from the backend service. The manager increments the file's version in the backend on each successful upload. For download, the manager compares the version against the locally cached version. If there is an update, the new file is downloaded in the background and overwrites the local cache and notifies all features using the file inside the app.
ishandotpage 6 hours ago [-]
I have had a particularly poor experience with Cloudflare Billing support, where an unexplained invoice was raised, silently? failed, and R2 access was immediately cut off, support was unhelpful in being able to explain anything AND took weeks to reply, and I had to move off R2 posthaste as it brought down production.
I will not be trusting Cloudflare with anything production critical.
scrubs 20 hours ago [-]
I forgot about R2. And I just got a problem perfect for it. Thanks for the reminder. HN helpful yet again!
elendilm 8 hours ago [-]
Glad I was helpful.
Any technology that lowers latency, memory, or cost must be appreciated and recommended.
AznHisoka 21 hours ago [-]
Is there a catch? So you can host a 1TB file and if 9,999,999 people download it you pay $0?
Dylan16807 21 hours ago [-]
There's no catch. They don't care about the bandwidth. If 20 million people make a single request each then you'll pay $3.60 to cover the requests.
And you're going to have to pay $15 of storage costs for your terabyte.
Cloudflare can shut you down and charge an arbitrarily high price if they don't like you and feel you're abusing their low prices.
kalleboo 18 hours ago [-]
The tl;dr for anyone skimming the comments: "CloudFlare doesn't like you" in this case means a gambling site that was getting CloudFlare shared IP addresses legally blocked by ISPs in countries where the gambling site wasn't licensed/was illegal. And "abusing low prices" means CloudFlare wanted them to pay for a bring-your-own-IP plan to give them dedicated IPs so their blocks wouldn't affect other CloudFlare customers.
Dylan16807 17 hours ago [-]
Cloudflare was really unclear about their motivation was and it was effectively a matter of kicking them off for not liking them. CF was annoying them about an overpriced enterprise plan, before suddenly throwing out accusations that if true would have meant cutting the company off even if they did buy enterprise. And they were demanding a ton of money, much more than bringing an IP is worth.
Maybe maybe bringing their own IP would have solved the problem, but Cloudflare was obstructing everything and then did a cutoff without proper warning. It was really bad on Cloudflare's part.
someonebaggy 15 hours ago [-]
This the same cloudflare that gets blocked in the entire country of Spain during every football match and says nothing about it?
They really hated that particular customer.
PunchyHamster 24 hours ago [-]
> S3’s dominance is due to its many advantages: effectively infinite capacity, high durability, and low per-gigabyte capacity cost.
It's not "cheap". 6 months of S3 is at around price of outright buying 4TB SSD at retail price. That before you do any IOPS to it
S3 is terrible deal on any front. it's just easy
0xCMP 24 hours ago [-]
As someone with TBs of SSDs and HDDs laying around I am someone who agrees with your point, but this is not a good way to compare costs.
For one, this has no redundancy and doesn't factor the cost of the machine providing the storage. But also it's an upfront cost for storing 4TB. It would be implied you wouldn't literally store 4TB on the SSD because now you're out of capacity for more data. If you only have 2GB of data in a bucket then S3 is still magnitudes cheaper, faster, and more redundant than anything you could put together yourself because the cost is spread across all the users of the service.
Reality is there are so many times that S3 makes the most sense that it makes people short-circuit and always pick S3 despite the huge hidden IOPS and bandwidth costs everyone rightfully tries to point out.
someonebaggy 22 hours ago [-]
If you have only 2GB of data it's trivial, it hardly matters how you store it. You might as well hold a full copy in RAM on each server and synchronize it with your developer laptop every half hour in case they all go offline at once, for all it matters.
vel0city 21 hours ago [-]
To reliably hold a bit over 2GB of data in RAM on AWS, you'll probably need something like a t4g.medium, about $0.0336/hr on-demand. 730 hours in a month, so $24.528, but ideally you'll have like three or so instances so like $74/mo. That's before thinking about stuff like IP address costs, the EBS volumes underpinning those boxes, etc. And then managing all the syncing and what not.
Of course, that's list prices, you can get savings plans and RIs and discounts.
The cost of RAM per hour in AWS isn't cheap.
someonebaggy 20 hours ago [-]
Once again proving that AWS is expensive any way you slice it. Even in today's RAM crisis, an extra 2GB is what, $50? and you probably already have 2GB free since they don't make RAM modules that small.
vel0city 8 hours ago [-]
Just buying a few sticks of RAM isn't hosting data in that RAM indefinitely. Its not the same as what S3 (or similar services) gets you.
What is your datacenter costs for multiple different buildings (assuming colocation/office space)? How much are you going to pay for the electricity? How much are you paying for networking? How much is rent for each U of server? How much do you spend on payroll to have knowledgeable people to manage it 24 hours a day, 7 days a week? Suddenly its not just a one-time $50 payment.
Maybe you don't actually need all of that, just some slices of some of that whole stack of things. Probably true in many situations. But just looking at it as a couple of sticks of RAM is not a like-for-like comparison.
I'm not arguing the cloud is cheap. You can definitely go cheaper and choose your own cost optimizations when managing the hardware yourself, deciding what is actually a necessary cost and what isn't. I totally agree AWS overchages for a lot of what they do; they price their stuff at a premium because they know businesses will pay it. But hosting 2GB of stuff in RAM reliably and indefinitely to anywhere near the guarantees of S3 is going to be a lot more than just a single $50 payment.
acdha 23 hours ago [-]
> It's not "cheap". 6 months of S3 is at around price of outright buying 4TB SSD at retail price. That before you do any IOPS to it
This is badly misunderstanding the problem: S3 is a highly available service with geographic redundancy and a huge range of integrated features. Your comparison would need to be updated to include multiple running servers in addition to redundant storage, and the software stack implementing things like immutability, not to mention all of the security features.
That’s not to say that you can’t build equivalents for the parts you use but you either need massive scale or giving up features to do that. For example, if you can tolerate bitrot or long access times if hardware fails, you can definitely get a lower cost per terabyte.
The reason why most people don’t do that, even when they have scale, is that it adds cost and risk everywhere else if you need engineering/ops people working on storage. If you’re, say, the internet archive that might make sense—it’s quite literally why your organization exists—but most other places are going to see all of the integrated features that they don’t have to build and operate paying for the difference between S3 and physical media pricing. For example, if I want to process files as soon as they’re uploaded or have immutability, I can just turn that on rather than having to build more services.
ehe78qhe 23 hours ago [-]
Even if you factor that it, S3 is pretty expensive for most small to medium size data.
acdha 23 hours ago [-]
Try doing the math and ask why your time costs. You need to buy a lot of storage to pay for the engineering work and most places would prefer to spend that time and attention on the product rather than shaving a few percent off of the storage bill (especially since the savings will be negative for quite some time).
ehe78qhe 23 hours ago [-]
I have done the math as part of my job and often found that yes, it was literally worth the money to do our own storage in cases. Running a highly available, durable storage system is not particularly difficult. There was an era when it was a common skill for a sysadmin. In the current era there are off the shelf solutions for it, both open and proprietary.
The biggest reason people pay for S3 is because of data transfer cost to other cloud services they are using.
everfrustrated 8 hours ago [-]
>Running a highly available, durable storage system is not particularly difficult
Only someone who had never done this would say this.
acdha 10 hours ago [-]
Yes, I’ve also built those services. As I said before, it is totally possible to beat S3 on pricing but the margin narrows considerably when you factor in the full costs — e.g. think about how many places had replicated storage but not bit-rot protection or versioning and built that at the application layer instead — and the savings is often not big enough to make it your top engineering pick for the money & attention required.
Fundamentally you’re playing a commodity game and that means you need a consistent edge to be worth the fixed costs. If your fully-loaded advantage is less than, say, 30% you’re better off negotiating with your AWS rep and investing your engineering resources in something your customers want.
aeonik 23 hours ago [-]
Wait, I'm confused, are you arguing for or against S3? Did you mean factoring the time needed to obtain the multiple PhD levels of information in tracking the cost, usage, and security of those interoperable services plugged into S3?
acdha 10 hours ago [-]
I’m saying that once you make the fully-loaded comparison, the gap closes a lot and most places will choose to spend their time and attention elsewhere unless they have a truly massive storage need.
Dylan16807 16 hours ago [-]
> a few percent
Really? No, it's not a few percent.
acdha 9 hours ago [-]
Again, spec out a multi-data center server + storage + staffing deployment and calculate the full cost per gigabyte. As I said earlier, if you buy enough storage you will definitely hit a level where you can beat S3 by a fair margin but many places never reach that point and many of the ones who do are going to have other projects which make more sense because they offer a high rate of return or provide something which isn’t a commodity.
Dylan16807 7 hours ago [-]
That's a good argument that at a lower amount of storage you won't save money at all.
This is very different from the idea that the savings are a few percent. In almost all cases the savings are nothing or a large fraction.
elendilm 23 hours ago [-]
Many of us just don't buy your argument. "Engineering work" for maintaining our own storage is not rocket science.
You just plug it in and basically it runs. Thats pretty much it.
Having to explain to an HN audience how installing an SSD is trivial is weird.
acdha 10 hours ago [-]
Again, this is completely misunderstanding the problem. Your single SSD with no redundancy, integrity, or security, or running servers competes with S3 in the same way that my backyard garden competes with local restaurants.
elendilm 9 hours ago [-]
You are shifting the goalpost.
If you don't pay the premium for redundancy, your S3 is the same sitting duck as the single SSD in your office, albeit an expensive one. Thinking AWS and therefore integrity and security is magical thinking.
Ask those who lost their data in Bahrain AWS.
But you are entitled to your expensive opinion. Some of us just don't buy it is all.
acdha 9 hours ago [-]
> You are shifting the goal-post.
No, that was literally the first paragraph of my original comment.
> If you don't pay the premium for redundancy, your S3 is the same sitting duck as the single SSD in your office, albeit an expensive one.
Again, this is misunderstanding what the service provides relative to raw storage. I suggest reading the documentation about reliability and thinking about how that SSD in your office fails to provide either geographic redundancy (which all but the one-zone flavor of S3 has) or bitrot protection. Similarly, while it is true that some security concerns still require you to configure things correctly, you should think about the infrastructure you would need to build to match the core service: none of that is something you can’t match but building and operating that has costs.
> Ask those who lost their data in Bahrain AWS.
Let’s think about this slightly longer: how is the SSD in your Bahrain office doing? Iran has to hit multiple AWS data centers to lose your S3 data so to match that you need to build a system which syncs changes at multiple separate campuses in the region, which is significantly more expensive, and when you realize that you probably want a copy somewhere far away, the work to build that will be more than checking a box — and don’t forget that you have to regularly test all of this to make sure you can trust it to work in a disaster.
Again, as readers of the original comment know, I’m not saying just you can’t do this but that it’s a category error to think the comparison is the raw hardware and not the combination of redundant hardware, software, and staffing you need to match capabilities. You can beat that with enough scale but that’s a very different equation.
elendilm 9 hours ago [-]
The original comment by PunchyHamster was about cost, i.e how AWS S3 base tier costs cannot be justified if you purchase your own SSD.
You shifted the conversation to redundancy and how AWS provides bells and whistles. You are consistent in your position, just not with me and not with PunchyHamster because we both are comparing costs by self hosting.
And for those of you who are willing to pay the premium and don't care about costs, like you, would obviously find AWS attractive.
For a startup like ours and many others, cost is everything. And spending a day or two for bitrot protection and having a couple of SSDs for redundancy and scalability is a no brainer.
You don't need big buildings for resilience. I suggest building things that require real scaling and resilience and you would know that it is about being distributed, having low cost, and owning your systems.
Your argument that it is expensive and tedious to maintain compared to AWS S3 is false. No it isn't.
Fear mongering only works against those who haven't built reliable systems. Building an S3 equivalent with cheap providers for redundancy or having one's own small resilient distributed hardware with staffing most often dwarfs the horribly insane premium paid for AWS S3 services.
But this is only if you want the redundancy which is a different topic altogether.
If all you want is a single replica, which was what PunchyHamster originally commented on, the base price of what AWS S3 provides, which does not include redundancy, makes even lesser sense compared to having your own SSD.
someonebaggy 22 hours ago [-]
Do you need that? If you want your own replicated storage cluster at home, you can use Ceph. I've seen Ceph deployed to utilize the spare hard drive ports in server clusters that otherwise mostly do compute (transcoding), alongside memcached to utilize the RAM, etc.
I'd bet the majority of projects, but perhaps not the majority of traffic, can run adequately on a single server and tolerate a few hour maintenance window one weekend every month. Don't waste money on overkill.
PunchyHamster 9 hours ago [-]
> This is badly misunderstanding the problem: S3 is a highly available service with geographic redundancy
First off, for geo-redundancy you need to pay extra, you only get region, and with recent strait crisis amazon already showed that can be hit and be unrecoverable.
And you're badly misunderstanding the point; that it is fucking expensive from every angle.
You can get geo-redundancy cheaper (their competitions prove it all the time); you can get IOPS cheaper; and with massive difference in end of the month bill.
Especially if you need a tons of bandwidth, S3 is very, very, very expensive way to get there.
It might be acceptable if you just run a website, but once you start running data-heavy operations on it, a big SSD attached to server can be order of magnitude cheaper than running it on S3 (and you can still use it to backup the data in case something fails)
Of course, if you had paid extra for the backups and redundancy, your data would survive.
So to address your original point, unless you are paying extra for redundancy, it is cheaper if you use a 4TB SSD at retail price as the original commenter discussed.
someonebaggy 22 hours ago [-]
In this case AWS did not save you. AWS is supposed to be resilient against AZ failures within a region. Iran destroyed all AZs at once, so all the data was lost. If you had a hard drive in your office, either directly serving your project or as an off-site backup, you still have your data. In the end, was it worth paying several times as much for AWS to provide durability for you, only to have them lose the data anyway?
everfrustrated 8 hours ago [-]
S3 has never been advertised as a DR strategy
elendilm 7 hours ago [-]
Focus on costs. That is the whole argument.
elendilm 21 hours ago [-]
You seem to have misunderstood. What you said is exactly my argument. I am with you :).
The parent commentator is under the illusion that AWS automatically means security and scale and reliability.
I was merely pointing out to him that Iranian attacks must serve as a wake up call for him.
inkyoto 19 hours ago [-]
[flagged]
hadlock 20 hours ago [-]
S3 cost growth is so gradual most businesses will just absorb the cost. We have some physical backups, but everyone is more than happy to not be shuffling around physical hard drives, and pay Amazon to deal with that and securely store it.
tempay 22 hours ago [-]
If you play the game correctly it can be a good deal. The ~1 USD per TB per month of glacier deep archive is hard to compete with if you (probably) don’t need to read the data back.
someonebaggy 22 hours ago [-]
I agree, but that's cherry-picking what is perhaps the only cost-effective service in all of AWS.
elendilm 22 hours ago [-]
What use is data that is never read unless it is for redundancy, disaster recovery, legal, or audit purposes, etc.
Yes you can make the argument for that specific use case.
For everyday use case, using s3 as your primary storage is costly and is not at all ideal.
CodesInChaos 24 hours ago [-]
Everything in AWS is expensive. But at least S3 adds a lot of value.
someonebaggy 19 hours ago [-]
Ever tried Ceph? It's a bit painful to set up and operate, but it seems to work pretty well.
(Don't bother with Ceph Object Gateway unless you really really need it - since you're programming your own application, access the storage pool directly with librados)
bigstrat2003 18 hours ago [-]
That's AWS in general for you. It has its place, but it's vastly overused in our industry. A lot of businesses would find their TCO would go down if they ditched cloud services, but that's not trendy enough for execs to do it.
Amazon S3 storage usage is calculated in binary gigabytes (GB), where 1 GB is 2^30 bytes. This unit of measurement is also known as a gibibyte (GiB), defined by the International Electrotechnical Commission (IEC). Similarly, 1 TB is 2^40 bytes, i.e. 1024 GBs.
I guess they could just change that overnight in the future. Still, a lot better than charging for Storage Units or similarly arbitrary unit.
[1]: https://aws.amazon.com/s3/pricing/#S3_Pricing_Details
GiB is unambiguously the big one, power of two. You want S3 to be priced in dollars per GiB. "shrinkflation power of ten GiB" is getting it backwards.
God why am I defending Amazon?
For reference transit cost has dropped over the years and is now about at $0.000247/GB or free if there is a peering agreement with the network the data is sent to.
This isn’t just AWS BTW; all tier 1 cloud providers recoup their costs this way.
And kindly refrain from accusing others of Stockholm Syndrome here. It’s incredibly rude.
Finding any excuse to justify a 100x profit margin is a form of Stockholm syndrome.
We are comparing to T1 ISPs because they also provide huge amounts of bandwidth reliably, and they do it 3 orders of magnitude cheaper. Therefore it cannot possibly be the case that providing the bandwidth is simply that expensive.
Many of us are aware that ISPs charge less for bandwidth. But they are not offering the same thing as a tier 1 cloud provider.
You may find these interesting:
https://aws.amazon.com/video/watch/c37546e1558/
https://cloud.google.com/blog/products/networking/speed-scal...
Again, these egress fees are not paying for egress alone. They are defraying the cost of operating a massive, complicated, low-latency, reliable cloud network infrastructure, much of which is free of charge when used internally.
Some, like you, are expecting that cloud providers operate on a “cost plus” model where what you pay is strictly based on what marginal costs for the same thing. But it’s not that simple in reality. Charges levied for one thing can be used to pay for another.
If it's surplus it's not costs. Why did you claim that the high price is to cover the costs if 99.9% of it isn't covering the costs?
If I may take the opposite view: companies should be grateful to get any profit from me at all.
I edited my comment above to explain in more detail what these charges are covering (and it’s not egress bandwidth alone). This has been a fast moving conversation and I’ve tried to add more context, unfortunately after the fact. My bad for posting too quickly.
They could charge 10x-ish as much as T1 ISPs and cover all those costs and still have it be almost all profit.
Show us the math.
It's no different than if I'm selling apples for $1000 each. Even if I can find one or two customers willing to pay that, wouldn't I have to be pretty stupid to spend $900 growing each apple incompetently when I could just spend $2 at the grocery store to fulfil those orders?
Yes, you're legally correct, but the point is AWS are clearly overcharging for egress.
As you say:
> Egress pricing pays for the AWS network infrastructure that is incredibly reliable (hardware failures happen regularly yet almost no one notices) and allows it to operate at scale sufficient to absorb even the largest DDoS attacks.
Cloudflare is also pretty good at absorbing DDoS attacks, yet charge no egress costs.
Maybe the real question is, are the tier-1 providers colluding in overcharging for egress?
Egress costs are fairly clearly a vendor lock in strategy. If I want to transfer 1TB to run on cheaper compute at another provider, that costs me $90 (edit) on AWS as opposed to $0.24 at wholesale prices. It's unlikely to be cost effective so I choose AWS compute.
EDIT: thanks to someonebaggy for pointing out my decimal point error.
Ninety dollars to transfer a terabyte is clearly ridiculous.
Many people avoid T1 clouds because they're so expensive, but they seem to have found an extremely profitable market segmentation consisting of the remainder.
There are a lot of very competitive alternatives around today are largely unknown. For example B2 Overdrive by Backblaze boasting up to 1 Tbps, Wasabi offering byte-compatible S3 storage with no egress fees for reasonable use and Cloudflare R2 also with no egress, all priced comfortably under S3 before egress (which the the real kicker...how do you even put a cost on lock-in?).
Not shilling any of these. We use R2 and Wasabi as our storage backends...so we obviously like them, but are not paid to say good things about them.
(1) https://www.backblaze.com/cloud-storage/b2-overdrive (2) https://wasabi.com/cloud-object-storage/wasabi-fire (3) https://carolinacloud.substack.com/p/pushing-r2-to-its-limit...
There's only a small little mention, if we are lucky, on the couple drives that have it (expensive enterprise flagships). It should be a regular sticking point, whether it's there or not. Without pressure it's not going to get regularly available, it feels like.
FDP is so simple. Declare a number for what pool of data you want to write into. Data of the same pool gets written to the same storage such that you can wipe it latter together. It has huge wins though against write amplification! Massive wins. For so close to free.
Some day I want to own a FDP drive. And then I can finally start using the tokio/io-uring support that I contributed! https://github.com/tokio-rs/io-uring/issues/380
The NVMe-KV is more radical. Still worth putting some pressure on, but your drive as KV, as object store, feels harder. Side note, really enjoyed this ceph nvme-kv offload post thing, my favorite tech write up in a while! https://ceph.io/en/news/blog/2026/for-whom-the-door-bell-tol...
FDP is a very minimal addition of control. But yes it is another box to ticket, is another place to upcharge. Yet still, we're only just seeing mainstream products emerge. Kioxia's CM10 for example. https://www.techpowerup.com/351218/kioxia-introduces-first-p...
I'm hoping that the need is great enough to break the industry control. I think for a while the market felt relatively well enough served such that it was unclear whether clear wins would actually result in customers. With AI need for speed, I can definitely imagine incredibly crazy CXL controllers or what not, that allow low latency acces to many many open channel flash systems. Wouldn't that be a thing.
Again though, FDP is such a ridiculously tiny add, and it helps SSDs so much for so many use cases. I really hope it becomes an expectation, not a feature, sooner rather than latter.
Edit: happy to see a new group has shown up asking for open channel flash, Open Flash Platform. https://openflashplatform.org/
The standards say SFTP. Most everyone ignores that part and have been using S3 buckets by bilateral consensus instead for years.
It turned out that all of our peers supported blob storage better than SFTP, which has some “show stopper” problems like forced outages caused by mandatory host key rotations.
But also while it should be faster than it is, how would an S3 designed around SSDs look differently to an API user? I would think the API is basically the same.
- Directory entries are no longer returned in sorted order in ListObjectsV2
- There's an AppendObject API
- There's a RenameObject API
I suspect it's more to do with the fact that with One Zone is a clean rewrite of large parts of the application stack that makes up S3.
S3 is made up of hundreds of microservices[0], there probably isn't anyone at Amazon that actually understands the whole system. Refactoring it to support these features probably requires coordination between a lot of different teams. They might have petabytes of metadata, making a change to how metadata is persisted probably requires a massive risky data migration.
[0]: "All in, S3 today is composed of hundreds of microservices" - https://www.allthingsdistributed.com/2023/07/building-and-op...
I’d also like a “shitter” storage service. I.e no redundancy, no multi-AZ. Just really cheap storage for stuff I don’t really care about losing, so that I can store a bunch of stuff in bulk that can easily be recovered.
What you need to understand is that a modern hard drives are severely limited in the amount of IOPS they can perform. Disk size has grown, but read and write speeds have not kept up with it.
In 1990 you could read a whole hard drive in 2 minutes, in 2026 as sizes have grown it takes the entire day.
As a system designer you want the maximum utilisation of your hardware, so you want to be close to saturating the IO that your drives are capable of.
What that means in practice is that only a tiny portion of the drive can be allocated to files that are read often.
That leaves the rest of the drive for files which are not used very often.
If you do a clever bit of caching work in your app like we have with our apps Slyp and SlypBusiness, you could even make it insanely cheap.
We have a file manager that is ultimately responsible for images, files, videos and whatnot. It loads the images/files from the downloaded local cache when requested by a feature in the app.
A standard feature uploads and downloads using signed urls obtained from the backend service. The manager increments the file's version in the backend on each successful upload. For download, the manager compares the version against the locally cached version. If there is an update, the new file is downloaded in the background and overwrites the local cache and notifies all features using the file inside the app.
I will not be trusting Cloudflare with anything production critical.
Any technology that lowers latency, memory, or cost must be appreciated and recommended.
And you're going to have to pay $15 of storage costs for your terabyte.
Cloudflare can shut you down and charge an arbitrarily high price if they don't like you and feel you're abusing their low prices.
Maybe maybe bringing their own IP would have solved the problem, but Cloudflare was obstructing everything and then did a cutoff without proper warning. It was really bad on Cloudflare's part.
They really hated that particular customer.
It's not "cheap". 6 months of S3 is at around price of outright buying 4TB SSD at retail price. That before you do any IOPS to it
S3 is terrible deal on any front. it's just easy
For one, this has no redundancy and doesn't factor the cost of the machine providing the storage. But also it's an upfront cost for storing 4TB. It would be implied you wouldn't literally store 4TB on the SSD because now you're out of capacity for more data. If you only have 2GB of data in a bucket then S3 is still magnitudes cheaper, faster, and more redundant than anything you could put together yourself because the cost is spread across all the users of the service.
Reality is there are so many times that S3 makes the most sense that it makes people short-circuit and always pick S3 despite the huge hidden IOPS and bandwidth costs everyone rightfully tries to point out.
Of course, that's list prices, you can get savings plans and RIs and discounts.
The cost of RAM per hour in AWS isn't cheap.
What is your datacenter costs for multiple different buildings (assuming colocation/office space)? How much are you going to pay for the electricity? How much are you paying for networking? How much is rent for each U of server? How much do you spend on payroll to have knowledgeable people to manage it 24 hours a day, 7 days a week? Suddenly its not just a one-time $50 payment.
Maybe you don't actually need all of that, just some slices of some of that whole stack of things. Probably true in many situations. But just looking at it as a couple of sticks of RAM is not a like-for-like comparison.
I'm not arguing the cloud is cheap. You can definitely go cheaper and choose your own cost optimizations when managing the hardware yourself, deciding what is actually a necessary cost and what isn't. I totally agree AWS overchages for a lot of what they do; they price their stuff at a premium because they know businesses will pay it. But hosting 2GB of stuff in RAM reliably and indefinitely to anywhere near the guarantees of S3 is going to be a lot more than just a single $50 payment.
This is badly misunderstanding the problem: S3 is a highly available service with geographic redundancy and a huge range of integrated features. Your comparison would need to be updated to include multiple running servers in addition to redundant storage, and the software stack implementing things like immutability, not to mention all of the security features.
That’s not to say that you can’t build equivalents for the parts you use but you either need massive scale or giving up features to do that. For example, if you can tolerate bitrot or long access times if hardware fails, you can definitely get a lower cost per terabyte.
The reason why most people don’t do that, even when they have scale, is that it adds cost and risk everywhere else if you need engineering/ops people working on storage. If you’re, say, the internet archive that might make sense—it’s quite literally why your organization exists—but most other places are going to see all of the integrated features that they don’t have to build and operate paying for the difference between S3 and physical media pricing. For example, if I want to process files as soon as they’re uploaded or have immutability, I can just turn that on rather than having to build more services.
The biggest reason people pay for S3 is because of data transfer cost to other cloud services they are using.
Only someone who had never done this would say this.
Fundamentally you’re playing a commodity game and that means you need a consistent edge to be worth the fixed costs. If your fully-loaded advantage is less than, say, 30% you’re better off negotiating with your AWS rep and investing your engineering resources in something your customers want.
Really? No, it's not a few percent.
This is very different from the idea that the savings are a few percent. In almost all cases the savings are nothing or a large fraction.
You just plug it in and basically it runs. Thats pretty much it.
Having to explain to an HN audience how installing an SSD is trivial is weird.
If you don't pay the premium for redundancy, your S3 is the same sitting duck as the single SSD in your office, albeit an expensive one. Thinking AWS and therefore integrity and security is magical thinking.
Ask those who lost their data in Bahrain AWS.
But you are entitled to your expensive opinion. Some of us just don't buy it is all.
No, that was literally the first paragraph of my original comment.
> If you don't pay the premium for redundancy, your S3 is the same sitting duck as the single SSD in your office, albeit an expensive one.
Again, this is misunderstanding what the service provides relative to raw storage. I suggest reading the documentation about reliability and thinking about how that SSD in your office fails to provide either geographic redundancy (which all but the one-zone flavor of S3 has) or bitrot protection. Similarly, while it is true that some security concerns still require you to configure things correctly, you should think about the infrastructure you would need to build to match the core service: none of that is something you can’t match but building and operating that has costs.
> Ask those who lost their data in Bahrain AWS.
Let’s think about this slightly longer: how is the SSD in your Bahrain office doing? Iran has to hit multiple AWS data centers to lose your S3 data so to match that you need to build a system which syncs changes at multiple separate campuses in the region, which is significantly more expensive, and when you realize that you probably want a copy somewhere far away, the work to build that will be more than checking a box — and don’t forget that you have to regularly test all of this to make sure you can trust it to work in a disaster.
Again, as readers of the original comment know, I’m not saying just you can’t do this but that it’s a category error to think the comparison is the raw hardware and not the combination of redundant hardware, software, and staffing you need to match capabilities. You can beat that with enough scale but that’s a very different equation.
You shifted the conversation to redundancy and how AWS provides bells and whistles. You are consistent in your position, just not with me and not with PunchyHamster because we both are comparing costs by self hosting.
And for those of you who are willing to pay the premium and don't care about costs, like you, would obviously find AWS attractive.
For a startup like ours and many others, cost is everything. And spending a day or two for bitrot protection and having a couple of SSDs for redundancy and scalability is a no brainer.
You don't need big buildings for resilience. I suggest building things that require real scaling and resilience and you would know that it is about being distributed, having low cost, and owning your systems.
Your argument that it is expensive and tedious to maintain compared to AWS S3 is false. No it isn't.
Fear mongering only works against those who haven't built reliable systems. Building an S3 equivalent with cheap providers for redundancy or having one's own small resilient distributed hardware with staffing most often dwarfs the horribly insane premium paid for AWS S3 services.
But this is only if you want the redundancy which is a different topic altogether.
If all you want is a single replica, which was what PunchyHamster originally commented on, the base price of what AWS S3 provides, which does not include redundancy, makes even lesser sense compared to having your own SSD.
I'd bet the majority of projects, but perhaps not the majority of traffic, can run adequately on a single server and tolerate a few hour maintenance window one weekend every month. Don't waste money on overkill.
First off, for geo-redundancy you need to pay extra, you only get region, and with recent strait crisis amazon already showed that can be hit and be unrecoverable.
And you're badly misunderstanding the point; that it is fucking expensive from every angle.
You can get geo-redundancy cheaper (their competitions prove it all the time); you can get IOPS cheaper; and with massive difference in end of the month bill.
Especially if you need a tons of bandwidth, S3 is very, very, very expensive way to get there.
It might be acceptable if you just run a website, but once you start running data-heavy operations on it, a big SSD attached to server can be order of magnitude cheaper than running it on S3 (and you can still use it to backup the data in case something fails)
Of course, if you had paid extra for the backups and redundancy, your data would survive.
So to address your original point, unless you are paying extra for redundancy, it is cheaper if you use a 4TB SSD at retail price as the original commenter discussed.
The parent commentator is under the illusion that AWS automatically means security and scale and reliability.
I was merely pointing out to him that Iranian attacks must serve as a wake up call for him.
Yes you can make the argument for that specific use case.
For everyday use case, using s3 as your primary storage is costly and is not at all ideal.
(Don't bother with Ceph Object Gateway unless you really really need it - since you're programming your own application, access the storage pool directly with librados)