NHacker Next
  • new
  • past
  • show
  • ask
  • show
  • jobs
  • submit
▲Postgres SELECT DISTINCT Does Not Scale (dbos.dev)
nattaylor 4 hours ago [-]
Tepix 1 hours ago [-]
As is mentioned in the article (now).
Dylan16807 59 minutes ago [-]
(since it was originally posted)
2 hours ago [-]
procaryote 25 minutes ago [-]
I've generally started to treat use of SELECT DISTINCT as a warning flag, as it's very common that it indicates bad code

Some use it because they don't understand uniqueness constraints and try to fix it in post so to say. Some use it because they forgot a join condition and are absolute amateurs. Some use it because it fixed a problem for them once and now they add it everywhere

These people seem to outnumber the people who use SELECT DISTINCT in a well thought out manner

Dylan16807 19 minutes ago [-]
It's definitely a warning flag if you're applying it to full-ish rows. This situation seems much more innocuous to me.
sandeepkd 1 hours ago [-]
It says that the company is co-founded by Postgres creator. I find that bit hard to believe given that there is nothing novel in the article, probably discovery for them. I do understand that everyone has to go through their own journey to learn these things but at the same time when you are running business then seeking professional help isnt a bad idea.

Based on my experience queries like these cannot scale, whatever you do. However if you are already on a path where you had invested a lot in such queries then hire a DBA, if you are not far off then hire an architect to model the data for better performance.

Dylan16807 1 hours ago [-]
Scale with what? If you have m distinct values in an index, then listing them this way takes m log(n) time, which is fine for many use cases no matter how much data you have.
sandeepkd 53 minutes ago [-]
The way the OP is trying to achieve all the goals by pushing the complexity on the queries/database is what I am referring to as non-scalable as data grows on SQL DB.

> If you have m distinct values in an index, then listing them this way takes m log(n) time, which is fine for many use cases no matter how much data you have.

And NO the runtimes are not right away applicable on machines at scale. You are dealing with DB locks, page sizes, available memory, existing data in memory, queue depth. Experienced folks get paid to short circuit such learnings

adrianN 33 minutes ago [-]
Runtimes are usually pretty well applicable at scale, it’s just that most people don’t have a good intuition about asymptomatic notation. Constants and lower order terms matter a lot in practice but are hidden in asymptotic notation.
Dylan16807 22 minutes ago [-]
Selecting distinct values of a single column with a simple condition is hardly pushing complexity into the database.
stemchar 60 minutes ago [-]
It's Stonebraker, he has a history of doing this to sell shit to people who don't need it.
colenikol2 4 minutes ago [-]
[dead]
thesuperevil 2 hours ago [-]
[flagged]
DiabloD3 5 hours ago [-]
"Postgres SELECT DISTINCT Does Not Scale"

Correct. This is documented in depth: DISTINCT sorts the results first.

The article's use case seems to imply the author did not know about GROUP BY, nor does it imply the author knew about indexes, nor ANALYZE. Postgres 18's new skip scan indexing also could help here, so ensuring the planner chooses that could help.

Dylan16807 5 hours ago [-]
Would GROUP BY fix the issue?

The article explains that skip scan doesn't do anything here.

> nor does it imply the author knew about indexes, nor ANALYZE

Indexes were talked about a lot, and they explicitly mentioned looking at the query plan.

DiabloD3 2 hours ago [-]
The article seems to have changed since I commented.
Dylan16807 1 hours ago [-]
Twisell 1 hours ago [-]
Why don't he use GROUP BY on indexed columns was my first thought also.

I guess we just have to patiently wait for the OP to hopefully read Postgres documentation from postgreSQL 10.x before correcting the article again.

Dylan16807 1 hours ago [-]
Does that fix it? Why would people be trying to add new index scanning modes if that's enough to fix it?

Someone else said it doesn't.

gfody 2 hours ago [-]
nope - either gets you a HashAggregate, GroupAggregate, or Unique depending on whether the column/input is indexed/ordered
tpetry 3 hours ago [-]
Did you even read the article? They show that a perfect index for their query didnt help because Skip Scan is currently not used for DISTINCT queries.
DiabloD3 2 hours ago [-]
The article seems to have changed since I commented.
joinHNtheysaid 1 hours ago [-]
[dead]
Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact
Rendered at 07:35:52 GMT+0000 (Coordinated Universal Time) with Vercel.