A blog or YouTube channel would be a nice halfway step to test the waters. A publisher might like to see that anyway. It would give you a natural feel for what the appetite is and where you want to put the level of the presentation.
ssivark 4 hours ago [-]
My personal opinion is that statistics textbooks usually come from a prescriptive perspective, and that makes it challenging for the reader to get visceral intuition for what is actually going on. Any reader would be far better off just visualizing the damn distribution / samples and using reasonable judgement, instead of implicitly assuming a Gaussians distribution and blindly memorizing tests / formulae. Making the distributions explicit allows us to model them and get an intuition for what the samples are telling us. I would whole-heartedly recommend the Model based machine learning book to anyone (online version is free) https://mbmlbook.com/
discardable_dan 15 minutes ago [-]
My larger issue, any time I have tried to learn statistics, is how fast the notation moves. You end up flipping back pages and pages just to double-check a definition that was given once and is now being extended syntactically. It's infuriating.
eru 3 hours ago [-]
How do you make 'reasonable judgements'? How do you tell whether someone else made reasonable judgements? How do you judge other people's intuition?
Look at the histogram and think about what distribution one could reasonably impute from samples. And what you would set as bounds for "outliers", per your needs. While we're at it, let me also say that it might be useful to specify outlier bounds not just based on the spread in sample values, but the costs/payoffs they imply for your application.
If you are not doing something crazy, most reasonable people would agree with your judgement. Conversely, if you are making non-obvious inferences where reasonable people disagree, you are in murky water and no sophisticated statistical method will save you. Math is not magic; theorems merely recycle (launder) modeling assumptions into results.
jbs789 45 minutes ago [-]
Yup. And thinking through the physical realities or whatever real world constraints exist.
stackghost 46 minutes ago [-]
>Look at the histogram and think about what distribution one could reasonably impute from samples. And what you would set as bounds for "outliers", per your needs. While we're at it, let me also say that it might be useful to specify outlier bounds not just based on the spread in sample values, but the costs/payoffs they imply for your application.
Only people with prior education/training in statistics are capable of doing this. The people who don't need a textbook.
Something like 60% of US adults read at or below the 6th grade level, and 25% of US adults struggle to comprehend graphs or charts entirely. Someone who has no idea what a standard deviation is can't intuit about distributions. I think you're dramatically overestimating the average person.
ssivark 37 minutes ago [-]
> Someone who has no idea what a standard deviation is can't intuit about distributions.
I disagree vehemently with this claim. I could cite my experience in teaching this topic to liberal arts / humanities college students in the US, but it is really more obvious than that. Anyone can understand a histogram easily and far more intuitively than they can understand the formula for a standard deviation and whether it must divide by N or N-1. Statistics courses and textbooks get stuck on that kind of pedantry, and most students end up missing the forest for the trees.
gus_massa 12 hours ago [-]
No idea about statistics, but in most physict courses in my university, they recomend 3 books:
1) The main book, that has a complete explanation and is well ordered. It's for learning.
2) Tha Landau book, that is super short and hard. It's only to check you didn't miss any important formula or topic.
3) There Feynman book, that is anassorted colection of fairytales for physicist. It's a pleasure to read it but you must already read 1 to understand it.
4) The Shaum book, that is almost a colection of exercices. Some people hate it. Some people love it. I like it as a companion to theother books.
I guess you are complaining that 1 is boring and want to write 3. It's a good idea, but it's harder than expected.
msla 6 hours ago [-]
Similar to an old idea I had about how every programming language needs three books:
1. Basic introduction.
2. Reference tome, which has absolutely everything.
3. Cookbook with style advice for the more advanced student, which assumes you've read 1 and can look up various details in 2.
These days, 2 would be a wiki and 1 would likely be a bunch of pages on that wiki, but it's still good if you have someone sit down and write 3.
WCSTombs 46 minutes ago [-]
I think you're describing the the Diátaxis framework [1], which would further split your (1) into fully guided tutorials and discursive explanations.
Back in my day, the 3 books for programmers were Knuth vol 1, Knuth vol 2, and Knuth vol 3. ;)
lukasbm 1 hours ago [-]
Most german speakers will look back at the 2000 pages of "Java ist auch eine Insel" in terror, but it was actually all three books in one.
actualeff0rt 1 hours ago [-]
I've got a Bachelors and two Masters degrees in CS/Math, but yet I feel Probability and Statistics is my greatest weakness. I just cannot grok it / build an intuition for it, and believe me, I've tried. My biggest gripe with Prob/Stats textbooks is that it's very hard to explain things without needing to rely on measure theory.
Maybe probability and statistics are a skill issue on my behalf, but what I absolutely loathe is the absolute lack of standardisation when it comes to notation in measure theory. Every textbook does it differently. All of them assume that their notation is the one everyone uses. Nobody bothers to explain _what_ the notation means. If you ask me, every bit of new notation should be introduced with a sentence or two on "how to read this symbol in your head" - especially when there are indices, subscripts and superscripts involved. It's especially terrible for measure theory because there's so much "implicit" information you're supposed to gather from the context - but in a way I understand it, because if every bit of notation of absolute and complete, I imagine it would be quite hard to type up.
Anyways, my rant on measure theory notation aside - I would absolutely read yet another Prob/Stats textbook. But unfortunately I will also drop it really quickly if the author doesn't show me any "notation-sympathy" :)
a_bonobo 45 minutes ago [-]
Did you see that Andrew Gelman and others just published Bayesian Workflows? https://avehtari.github.io/Bayesian-Workflow/
It sounds similar to what you're after, away from describing the logic of models and the maths, instead it's about (quote from intro) 'There are all sorts of tacit knowledge in applied statistics that do not make it into published
papers and textbooks. The present book is intended to put some of these ideas out in the open'
Where would your book fit into this?
montalbano 3 hours ago [-]
How would it overlap or differ from 'Statistical Rethinking'?
This is widely regarded as the most accessible intro textbook to Bayesian statistics.
+ Krushke’s Doing Bayesian Data Analysis, and for the very basics there is Downey’s Think Bayes. I guess there might be a gap in the literature for a different approach, heck, I’d read it, but there is some very good material already out there.
geokon 1 hours ago [-]
Maybe my expectations were too high given all the online praise, but I've been working through it for the past two weeks and I've been underwhelmed:
- The language is often very vague and imprecise.
- Concepts are introduced at random and then not used till much later. So as you're reading you're left scratching your head as to why something was brought up.
- There are constant philosophical and historical digressions that maybe only hold some deeper meaning on a second reading once you already know the topic.
- Constant references to other people, like some guy Fisher, who don't like the method. This is discussed at length and they keep combing back to this theme (they seem really butthurt about this?). But the craziest part is this all done before you even really understand what the method is!!
- The editor must have placed some strict requirement of saying "Baysian" at least five times per page.
- No index. Useless table of context. But lots of end-notes you feel compelled to flip to constantly
Overall it feels like a textbook written to impress other statistics professors - and as an outlet of the author to air some frustrations with people statistics (which may be completely valid!)
The overall structure and objectives seem solid for the most part. It's just a lot of the details aren't great. The examples are fun and compelling, but you have to do your own legwork to actually pick through all the prose and tie the pieces together to figure how it fits together mathematically. Fortunately AI helps as a tutor
howunfortunate 4 hours ago [-]
I'm a nerd and do a lot of stats for my job and I would not read a statistics textbook.
I read a lot of informational things, but math / stats / software has always felt like an area where a book is just the wrong format.
If I were you I'd make an interactive website like SQLZoo or a video series like StatQuest.
Those are educational formats that really clicked with me for whatever reason.
mdspan 1 days ago [-]
Having also studied statistics in university (undergrad), something I kept running into is that you can't really unlock the intuition for many concepts without taking more advanced courses. For example, degrees of freedom shows up as early as AP Statistics, but even a non-rigorous visual explanation of it leans on linear algebra, which most students don't see until much later.
I think more resources like seeing-theory would be great since stats books are almost universally dry (Blitzstein being a notable exception), but I'm not sure how easily more advanced concepts lend themselves to visual explanation in a way that's digestible for a non-stats person.
usernametaken29 21 hours ago [-]
I feel this boils down my learning journey as well. You start unraveling a very good intuition about the underlying concepts MUCH much later, but partly because those intuitions themselves are never conveyed and are supposed to be learned from the proofs, and are an indirect product of learning.
jldugger 2 hours ago [-]
A long time ago, I took a "statistics for engineers" class in order to graduate. I slept through most of the classes. It sucked, and 70 percent of it was just "distribution of the week." It did not help that homework was optional for half of it.
At some point in my professional career I started reading non-fiction books and even bought a used statistics textbook for 10 bucks on abebooks. I didn't end up actually reading it until 12 years later during the COVID lockdown. I ended up shooting for 10 pages a day, 7 days a week. If those 10 pages included review exercises, it would be a long night.
Could just be the right book at the right time, but this one really helped me understand stuff beyond the normal HS math stuff, like RMS-error, calculating correlation, the difference between standard error and standard deviation, the relationship between sample size and standard error, t-tests, and chi-squared. Working as an SRE/release engineer, this stuff really helped me overcome a lot of _bad_ canary data analysis my predecessors had constructed.
That book was the 3rd edition of Statistics by Freedman et al.[1] One thing I want to complement was getting the pedagogy right. Most chapters have strong narrative hooks, several "check your knowledge" problems, review exercises, and post chapter bullet points to assist with spaced repetition. There's even a series of "special" review exercises covering entire sections of the book, ie exams.
For the HN crowd I should also probably note that the book is almost entirely non-bayesian and not intended to prepare readers for further coursework. You will not learn normal phraseology like "IID," "random variable" or "kernel".
Even if I would not read it back to back on release, it would be a pleasure to have a reliable and citable reference on hand. Whenever I stumble upon new complex problems, outside of the regular, often fairly repetitive statistical questions of my field, I need to rely on many searches and LLM queries to find my answer. I wonder if a book could actually replace all that, but it would be my first place to check.
jadermcs 2 hours ago [-]
Definitely there’s an interest for visual pedagogical content. However a book nowadays may not be the most effective way to reach a wider audience, instead of a video or an interactive website. I guess that combining these other media may help you reach more people to get interested in the book.
Don’t ask. Start from a few blog posts and see traction, read feedback.
You will also see hos long it takes - and what is thd difference between an idea and making it real.
RobGR 5 hours ago [-]
I would at least investigate it. I have purchased and read statistics books recently. Mostly the old classics by R.A. Fisher and D.R. Cox and etc, I have a copy of Handbook 91 by Mary Natrella. My questions are pretty simple. I like the "worked examples" approach in Handbook 91.
ludicrousdispla 3 hours ago [-]
I took my first statistics class as an evening course, before going to graduate school, and 90% of the work involved doing hand calculations, avg., variance, std. dev., z-scores, t-tests, etc. And I think that gave me a strong grasp of those fundamentals.
If you could do something similar for bayesian statistics I think that would be useful, but not necessarily popular.
clutter55561 2 hours ago [-]
Yes, I’d read it.
But beware of opinions.
Don’t let people put you down, especially here in HN, where people are perceived to smart. Smart doesn’t equal sensible or unbiased.
Many books are written to scratch the itch of the author. Just like an open source project. It is a work of love.
itake 2 hours ago [-]
I probably would buy/read it, but I wouldn't rush out to buy one. I barely passed stats in college and I wish had a better intuitive understanding.
3 hours ago [-]
throwaway81523 3 hours ago [-]
Aren't a zillion statistics books out there already? Yes I've been wanting to read one, and Wikipedia also has lots of good statistics articles. I've been wanting to work through Freedman and Pisani's book but you know how it goes. It's supposed to be excellent.
gignico 2 hours ago [-]
This sounded to me like that old “would you steal a car?” anti-piracy campaign until I clicked to read the post :)
junon 4 hours ago [-]
Not sure I'd read it as-written. But I'd love a statistics book whereby the chapters are real case studies. I personally learn best when I can apply new theory to a tangible problem (not just an example problem that's been reduced to almost nothing).
However going the 'visual' route might be enough for me to pick it up.
BoredomIsFun 2 hours ago [-]
Yes surely. I like ML, >D>S and such but lacking in stats background I wishh I had.
Yes, with the caveat that someone else with knowledge in the field has to recommend it to me. Partner up with someone who teaches statistics.
pessimizer 5 hours ago [-]
Such a bizarre question. Many, many people have read statistics textbooks.
Can you write one that's more worth reading than the standard ones? Don't answer that question, just prove it.
bigdict 4 hours ago [-]
I would! If you are thinking to write one, do it!
aghuang 1 days ago [-]
It depends on what parts of statistics is being taught and the application of each of the leanings and how it relates to the real world.
More generally, I would buy a statistics if it is linked to today's interesting technological breakthroughs and also if it comes as a distilled version for beginners.
sdcfgy 2 hours ago [-]
I’d probably sample it randomly yes.
Crap jokes aside, I mostly use Statistics In A Nutshell. It’s pretty ok. I say that as a member of the RSS for 30 odd years.
rootsudo 4 hours ago [-]
Yes, I love it and would do so
Saline9515 1 hours ago [-]
I would read a statistics textbook, if there were plenty of exercises, with the answer and explanation, which is usually lacking in textbooks.
Ready the theory is nice but you only learn by solving problems that uses it.
exe34 1 hours ago [-]
Hi, I would read it if it came with a GitHub of python code that I can modify for my own use. The fewest dependencies, the better.
I like Statistical Rethinking, but unfortunately the solver was very slow and I didn't end up using it much - not the author's fault, it's probably my computer that's way too old.
hollowturtle 3 hours ago [-]
I would read it
eimrine 2 hours ago [-]
I have some statistic book in paper, trying to solve some lemmas from time to time, so I will not read any slop.
gdulli 5 hours ago [-]
Probably.
whattheheckheck 3 hours ago [-]
Yes foundations of agnostic statistics and all of statistics
aspectmin 3 hours ago [-]
I’m weird. I read textbooks and instruction manuals. I read incredibly fast though, so it fits nicely. There’s always some deeper learning to be gained in those pages.
Good luck if you do this.
rramadass 4 hours ago [-]
Yes, there is always a need for another "intuitive statistics" book.
However the link you have provided is not the way; it is low on content and high on pretty distractions. Use all sorts of diagrams and graphs primarily, with animations only where required. The key is to always relate to something in the real world so one can see its actual relevance. Also tie it back to other fields of mathematics so one can see how they all come together.
And of course Nassim Taleb's works are a good source of inspiration. Here is a great video summarizing Taleb's ideas nicely Pareto, Power Laws, and Fat Tails - https://www.youtube.com/watch?v=Wcqt49dXtm8
PolGraciaSerrah 17 hours ago [-]
[flagged]
Rendered at 09:09:50 GMT+0000 (Coordinated Universal Time) with Vercel.
Modelling distributions explicitly sounds nice, yes.
If you are not doing something crazy, most reasonable people would agree with your judgement. Conversely, if you are making non-obvious inferences where reasonable people disagree, you are in murky water and no sophisticated statistical method will save you. Math is not magic; theorems merely recycle (launder) modeling assumptions into results.
Only people with prior education/training in statistics are capable of doing this. The people who don't need a textbook.
Something like 60% of US adults read at or below the 6th grade level, and 25% of US adults struggle to comprehend graphs or charts entirely. Someone who has no idea what a standard deviation is can't intuit about distributions. I think you're dramatically overestimating the average person.
I disagree vehemently with this claim. I could cite my experience in teaching this topic to liberal arts / humanities college students in the US, but it is really more obvious than that. Anyone can understand a histogram easily and far more intuitively than they can understand the formula for a standard deviation and whether it must divide by N or N-1. Statistics courses and textbooks get stuck on that kind of pedantry, and most students end up missing the forest for the trees.
1) The main book, that has a complete explanation and is well ordered. It's for learning.
2) Tha Landau book, that is super short and hard. It's only to check you didn't miss any important formula or topic.
3) There Feynman book, that is anassorted colection of fairytales for physicist. It's a pleasure to read it but you must already read 1 to understand it.
4) The Shaum book, that is almost a colection of exercices. Some people hate it. Some people love it. I like it as a companion to theother books.
I guess you are complaining that 1 is boring and want to write 3. It's a good idea, but it's harder than expected.
1. Basic introduction.
2. Reference tome, which has absolutely everything.
3. Cookbook with style advice for the more advanced student, which assumes you've read 1 and can look up various details in 2.
These days, 2 would be a wiki and 1 would likely be a bunch of pages on that wiki, but it's still good if you have someone sit down and write 3.
[1]: https://diataxis.fr/
Maybe probability and statistics are a skill issue on my behalf, but what I absolutely loathe is the absolute lack of standardisation when it comes to notation in measure theory. Every textbook does it differently. All of them assume that their notation is the one everyone uses. Nobody bothers to explain _what_ the notation means. If you ask me, every bit of new notation should be introduced with a sentence or two on "how to read this symbol in your head" - especially when there are indices, subscripts and superscripts involved. It's especially terrible for measure theory because there's so much "implicit" information you're supposed to gather from the context - but in a way I understand it, because if every bit of notation of absolute and complete, I imagine it would be quite hard to type up.
Anyways, my rant on measure theory notation aside - I would absolutely read yet another Prob/Stats textbook. But unfortunately I will also drop it really quickly if the author doesn't show me any "notation-sympathy" :)
Where would your book fit into this?
This is widely regarded as the most accessible intro textbook to Bayesian statistics.
https://xcelab.net/rm/
- The language is often very vague and imprecise.
- Concepts are introduced at random and then not used till much later. So as you're reading you're left scratching your head as to why something was brought up.
- There are constant philosophical and historical digressions that maybe only hold some deeper meaning on a second reading once you already know the topic.
- Constant references to other people, like some guy Fisher, who don't like the method. This is discussed at length and they keep combing back to this theme (they seem really butthurt about this?). But the craziest part is this all done before you even really understand what the method is!!
- The editor must have placed some strict requirement of saying "Baysian" at least five times per page.
- No index. Useless table of context. But lots of end-notes you feel compelled to flip to constantly
Overall it feels like a textbook written to impress other statistics professors - and as an outlet of the author to air some frustrations with people statistics (which may be completely valid!)
The overall structure and objectives seem solid for the most part. It's just a lot of the details aren't great. The examples are fun and compelling, but you have to do your own legwork to actually pick through all the prose and tie the pieces together to figure how it fits together mathematically. Fortunately AI helps as a tutor
I read a lot of informational things, but math / stats / software has always felt like an area where a book is just the wrong format.
If I were you I'd make an interactive website like SQLZoo or a video series like StatQuest.
Those are educational formats that really clicked with me for whatever reason.
I think more resources like seeing-theory would be great since stats books are almost universally dry (Blitzstein being a notable exception), but I'm not sure how easily more advanced concepts lend themselves to visual explanation in a way that's digestible for a non-stats person.
At some point in my professional career I started reading non-fiction books and even bought a used statistics textbook for 10 bucks on abebooks. I didn't end up actually reading it until 12 years later during the COVID lockdown. I ended up shooting for 10 pages a day, 7 days a week. If those 10 pages included review exercises, it would be a long night.
Could just be the right book at the right time, but this one really helped me understand stuff beyond the normal HS math stuff, like RMS-error, calculating correlation, the difference between standard error and standard deviation, the relationship between sample size and standard error, t-tests, and chi-squared. Working as an SRE/release engineer, this stuff really helped me overcome a lot of _bad_ canary data analysis my predecessors had constructed.
That book was the 3rd edition of Statistics by Freedman et al.[1] One thing I want to complement was getting the pedagogy right. Most chapters have strong narrative hooks, several "check your knowledge" problems, review exercises, and post chapter bullet points to assist with spaced repetition. There's even a series of "special" review exercises covering entire sections of the book, ie exams.
For the HN crowd I should also probably note that the book is almost entirely non-bayesian and not intended to prepare readers for further coursework. You will not learn normal phraseology like "IID," "random variable" or "kernel".
[1]: https://www.amazon.com/dp/B00SLB5Q72?lv=shuf&channelId=520&p...
Another exemple of a successful visual pedagogical content is: https://www.byhand.ai/
You will also see hos long it takes - and what is thd difference between an idea and making it real.
If you could do something similar for bayesian statistics I think that would be useful, but not necessarily popular.
But beware of opinions.
Don’t let people put you down, especially here in HN, where people are perceived to smart. Smart doesn’t equal sensible or unbiased.
Many books are written to scratch the itch of the author. Just like an open source project. It is a work of love.
However going the 'visual' route might be enough for me to pick it up.
Can you write one that's more worth reading than the standard ones? Don't answer that question, just prove it.
More generally, I would buy a statistics if it is linked to today's interesting technological breakthroughs and also if it comes as a distilled version for beginners.
Crap jokes aside, I mostly use Statistics In A Nutshell. It’s pretty ok. I say that as a member of the RSS for 30 odd years.
Ready the theory is nice but you only learn by solving problems that uses it.
I like Statistical Rethinking, but unfortunately the solver was very slow and I didn't end up using it much - not the author's fault, it's probably my computer that's way too old.
Good luck if you do this.
However the link you have provided is not the way; it is low on content and high on pretty distractions. Use all sorts of diagrams and graphs primarily, with animations only where required. The key is to always relate to something in the real world so one can see its actual relevance. Also tie it back to other fields of mathematics so one can see how they all come together.
A good example to study is How to Measure Anything: Finding the Value of Intangibles in Business by Douglas Hubbard. Detailed review at - https://www.lesswrong.com/posts/ybYBCK9D7MZCcdArB/how-to-mea...
And of course Nassim Taleb's works are a good source of inspiration. Here is a great video summarizing Taleb's ideas nicely Pareto, Power Laws, and Fat Tails - https://www.youtube.com/watch?v=Wcqt49dXtm8