Join 📚 Josh Beckman's Highlights
A batch of the best highlights from what Josh's read, .
In 2022, production-grade vector databases were relying on in-memory storage at $2+ per GB, not counting the extra cost for durable disk storage. This is the most expensive way to store data. You can improve this by moving to disk, with triply replicated SSDs at 50% storage utilization which will run you $0.6 per GB. But we can do even better by leveraging object storage (like S3 or GCS) at around $0.02 per GB, with SSD caching at $0.1 per GB for frequently accessed data. That’s up to 100x cheaper than memory for cold storage, and 6-20x cheaper for warm storage!
Turbopuffer: Fast Search on Object Storage
turbopuffer.com
If you’re debating a decision point and you find yourself in a stalemate, one really powerful question to use is “What would cause you to change your mind?”.
This forces the other party to put aside their attempts at convincing you of the merits of their position, and instead puts them in a position where they need to challenge their own perspective.
The act of challenging one's own position may lead to a realization that they were wrong.
Discovering this for one's self is far far more powerful than having someone else try to convince you of a different perspective.
Jeff Bruton Thoughts
Jeff Bruton
The generation of a cache key is always important, but even more so with GraphQL. The dynamic nature of GraphQL queries is such that even a white space in the query could affect the key and cause a miss, even though it was the same query in the first place. A good cache key should generally contain at least:
User information (if authenticated API).
A query hash, which should be normalized as much as possible.
The variables hash (we would not want queries with different variables to be cached as the same thing).
The operation name
A cache-busting element.
Production Ready GraphQL
Marc-Andre Giroux
...catch up on these, and many more highlights