This is anecdote, not proper feedback, since I wasn't directly involved in the topic.
My company relied on PG as its search engine and everything went well from POC to production. After a few years of production and new clients requiring volumes of data an order of magnitude above our comfort zone, things went south pretty fast.
Not many months later but many sweaty weeks of engineering after, we switched to ES and we're not looking back.
tl;dr; even with great DB engineers (which we had), I'd suggest that scale is a strong limiting factor on this feature.
Can you tell us about your scoring function? Selecting 40M results from a dataset of ~1B and returning the top 10 based on some trivial scoring function is easy, any reasonable search system will handle that. The problem is when you have to run your scoring function on all 40M matching docs to decide which 10 are the most relevant. It's even more of a problem when your scoring function captures some of the complexities of the real word, rather than something trivial like tf/idf or bm25.
My company relied on PG as its search engine and everything went well from POC to production. After a few years of production and new clients requiring volumes of data an order of magnitude above our comfort zone, things went south pretty fast.
Not many months later but many sweaty weeks of engineering after, we switched to ES and we're not looking back.
tl;dr; even with great DB engineers (which we had), I'd suggest that scale is a strong limiting factor on this feature.