Weird how the readme talks about hash collisions, but not in the way I would expect. Two blocked domains with the same hash isn't a problem, they both need to be blocked.
Where hash collisions matter is false positives, e.g. if hash(google.com) = hash(adserver.com). There does appear to be a web dashboard and /unblock api so should be straightforward to resolve.
Yeah, I think they have it backwards. As you say, regardless of how many entries you have locally, the collision risk comes from false positive hash matches not from worrying if the positive hashes collide.
> Two blocked domains with the same hash isn't a problem, they both need to be blocked.
This is a problem if the second one is actually a major useful domain. If HN and adserver have the same hash, you have a problem. This could be described as a false positive for HN.
> Where hash collisions matter is false positives, e.g. if hash(google.com) = hash(adserver.com)
This sounds like the same issue I describe above where you need to use Google but block adserver. Otherwise it's only a problem if you implemented allow lists, right? This is when you explicitly allow Google and implicitly allow adserver along with it because of the hash collision.
If you're ever thinking about getting an ESP32-C3, do yourself a favor and get the variant that you can connect an external antenna to. The normal C3 has a built-in antenna that is extremely weak, making it useless for most things I was planning using it for.
I'm a bit of a paranoid freak that likes lists, so I've got a PiHole that has a total list of 14-16M blocked domains.
Great idea, and good for casual blocking, but I'm almost moving to an "allow list" mindset. This solution would probably work better for that, I wonder if the good parts of the internet would fit into an 140k list.
I've aggregated the lists into four tiers: Base, Recommended, Aggressive, and Paranoid. And I've got two sets of these, one that includes Newly Registered Domains (Full directory), and one that doesn't include NRDs (NoNRD directory, because NRDs are ... heavy: Paranoid list with NRDs: 9.1 million records, without NRDs: 3.7 million records).
Aggregation of lists under the topic that matches the file name. Fake News list needs to be updated as it contains a _way_ overbroad list that I need to remove from the aggregation.
Depends on the user but not everyone is visiting new websites everyday
Even for those that are, the number of domain-IP mappings needed will be relatively small
Definitely under 140,000
Most DNS data I use is "static", it rarely changes. As such most times I don't have to make DNS queries. I store the domain-IP mappings in proxy memory; this is faster than DNS
This is all a guesstimate, but my gut feel is that a considerable part of the HN population doesn't even read the linked content, only the comment section here.
With github being the exception, I use comments to see if the article is AI. If it is, I prefer the condensed opinions in the comments. Not to say it can't be interesting I just don't want to waste time reading overly verbose generated text.
I think it could fit even more domains using a Bloom Filter or similar probabilistic data structure. With a chance of false-positives of course. But that's a trade-off the project already makes.
pinhole is great but if the device fails your network is down. I swapped to nextdns because of that. maybe support of multiple devices would fix it - now that is this cheap it would not matter
.
The simple answer is have 2. If the cost per unit is low enough than deploying 2 PiHoles is trivial.
The developers could make users' lives easier by implementing a clustering option that syncs the 2 out of the box. For static deployments, where you don't change the config too often, it's still decent as it is.
I did the same (until my RPi zero stopped working). Both the devices share same vertical IP address and if one ever went down, the other would take over without any delay.
The biggest benefit of my local dns server is latency.
On wired internet, my dns is <1ms from my PC.
Upstream dns for me is pretty quick. Google and cloudflare dns are ~5ms from me.
But WiFi latency alone is ~8ms most of the time in my experience. On my fiber internet, pinging a dns server in some random upstream server miles away is lower latency than WiFi 10 feet away. But the real issue on WiFi is any packet loss at all adding 50-100ms to that at random depending on interference.
With DNS you are paying this latency cost all the time on nearly every request.
Please just write the first sentence of your README yourself
I dislike when LLMs talk to me like that. I hate it when humans do it, confidently spewing their overconfident llm assumptions to others
Where hash collisions matter is false positives, e.g. if hash(google.com) = hash(adserver.com). There does appear to be a web dashboard and /unblock api so should be straightforward to resolve.
This is a problem if the second one is actually a major useful domain. If HN and adserver have the same hash, you have a problem. This could be described as a false positive for HN.
> Where hash collisions matter is false positives, e.g. if hash(google.com) = hash(adserver.com)
This sounds like the same issue I describe above where you need to use Google but block adserver. Otherwise it's only a problem if you implemented allow lists, right? This is when you explicitly allow Google and implicitly allow adserver along with it because of the hash collision.
Great idea, and good for casual blocking, but I'm almost moving to an "allow list" mindset. This solution would probably work better for that, I wonder if the good parts of the internet would fit into an 140k list.
Since it's not ready, it'll be best if I point you to the sources I've aggregated from:
https://firebog.net/
https://github.com/hagezi/dns-blocklists
https://github.com/StevenBlack/hosts
https://github.com/jerryn70/GoodbyeAds
Link to my in-progress list file hosting: https://lists.uninvitedactivity.com/DNS/
00_Tiers directory:
I've aggregated the lists into four tiers: Base, Recommended, Aggressive, and Paranoid. And I've got two sets of these, one that includes Newly Registered Domains (Full directory), and one that doesn't include NRDs (NoNRD directory, because NRDs are ... heavy: Paranoid list with NRDs: 9.1 million records, without NRDs: 3.7 million records).
11_Allow directory:
A bunch of allow lists sourced from: https://discourse.pi-hole.net/t/commonly-whitelisted-domains.... I also use this list for an Allow list: https://github.com/anudeepND/whitelist
21_SpecificTopics directory:
Aggregation of lists under the topic that matches the file name. Fake News list needs to be updated as it contains a _way_ overbroad list that I need to remove from the aggregation.
Depends on the user but not everyone is visiting new websites everyday
Even for those that are, the number of domain-IP mappings needed will be relatively small
Definitely under 140,000
Most DNS data I use is "static", it rarely changes. As such most times I don't have to make DNS queries. I store the domain-IP mappings in proxy memory; this is faster than DNS
No "blocklist" needed
If i would estimate it, I only take a look at 1/4 of the links where i read the comments.
The developers could make users' lives easier by implementing a clustering option that syncs the 2 out of the box. For static deployments, where you don't change the config too often, it's still decent as it is.
I would only recommend using it instead of the standard Arduino Processing IDE.
Upstream dns for me is pretty quick. Google and cloudflare dns are ~5ms from me. But WiFi latency alone is ~8ms most of the time in my experience. On my fiber internet, pinging a dns server in some random upstream server miles away is lower latency than WiFi 10 feet away. But the real issue on WiFi is any packet loss at all adding 50-100ms to that at random depending on interference.
With DNS you are paying this latency cost all the time on nearly every request.
Surely not? macOS, Windows, and I think most of the big Linux distros, all cache DNS responses. Probably the web browser itself does too.
Jokes aside, very impressive that it works!
after all, why use opnSense; just use suricata and $PROPER_X! why use TrueNAS; just use $DISTRO and ZFS!