When the Debugger Lies

(danielmangum.com)

48 points | by hasheddan 2 days ago

4 comments

  • pletnes 16 minutes ago
    Reminds me of the time when visual studio would show one value of a variable, but print would show a different one. Never trusted VS again after that.
  • Loren_SL 1 day ago
    Nice writeup. WinDbg has the same stale-cache trap on live Windows targets.
  • zavec 3 hours ago
    Very cool! I love reading about details like this.
  • LoganDark 3 hours ago
    Once upon a time I had to figure out an issue where I forgot a `return` statement at the end of a non-`void` function, so the C++ compiler happily omitted both RETs for some reason and let the program go straight into illegal instructions. That was fun to debug (not fun, I practically had to single-step through the entire program) because every time this happened the debugger was incredibly confused about what the fuck was going on and nothing made any sense.
    • VorpalWay 3 hours ago
      It makes no sense to me that C and C++ didn't require a diagnostic for this, rather you hit UB.

      I know that at least modern GCC and Clang will warn and/or error for this (not sure which as I use -Werror), but still, this is pointless UB to have.

      And no, in this case I don't buy that a C90 compiler would have been unable to check this.

      • jcranmer 2 hours ago
        Having been involved in the WG14 discussions on this topic:

        The issue is that there is a contingent of users who complains about cases where the return dynamically can't be hit but that isn't obvious statically. Consider something like this:

          int do_something(enum meow koala) {
            switch (koala) {
            case enum_val_1: return 5;
            case enum_val_2: return 3;
            /* etc., covering all the enum values */
            }
          }
        
        Should this be required to diagnose? That's the sticking point.
        • VorpalWay 1 hour ago
          There may be differences between C and C++ here (and I worked with C++ more recently than C), but: if that enum doesn't specify an underlying type it would be UB to have a value that isn't in the "member list" of the enum. In that case it should be possible to determine if all cases are covered.

          If an underlying type is specified, then it should error since it is legal to have those values (unless the whole range of the underlying type is covered by the cases.

          Again, that is what would be sensible from a C++ perspective, I don't know if C differs here.

          EDIT: Also, and now I'm talking with my Rust user hat on: it is better to not have pointless UB. Yes some is needed to practically allow for optimisation. But C and C++ had a lot of UB that doesn't really help with making your code faster, such as this.

          • hn_go_brrrrr 6 minutes ago
            I don't think your C++ comment is true, otherwise you wouldn't be allowed to OR together C enums. I think the restriction is on values wider than the underlying type, where the underlying type is always wide enough to support any representable bit pattern.
          • throw324523 1 hour ago
            > Again, that is what would be sensible from a C++ perspective, I don't know if C differs here.

            In C it is not UB for an enum to have an integer value that does not correspond to any listed enumeration constants.

      • LoganDark 2 hours ago
        I think I was using C++14 at the time...
      • Brian_K_White 2 hours ago
        It seems like something that could have been trapped right in k&r.

        Was it ever even theoretically under any circumstances for any reason intended to be able to write a stack of functions with no returns that just fall into each other like assembly? I can't believe it.

        So it seems like something even the very first compiler could have cought right in an early parser pass or stage.

        But I also decline to believe I have a better idea about something than K or R, so there must be a non-triviality I don't see. I mean goto() exists in the language so ?

        ... I guess simply detecting the end of a function, or detecting that the process reached the end of a function, isn't a good enough definition of the problem. You can have any number of returns or gotos in the middle that you are always supposed to hit, and intentionally no return at the end because instead you have an assert or a goto.

        assert you should never get here, goto error, goto not error but just next step, etc. They might or might not be error conditions that the process reached that spot, but it's not an error that the code doesn't end with a return.

    • StilesCrisis 3 hours ago
      Real
      • LoganDark 3 hours ago
        Another issue that happened in the same project is that for some reason whenever I compiled it with a regular C++ compiler, field writes were disappearing into the abyss. I think I actually never figured that out because it was literally the same object at the same memory address, there were no other threads and yet when a subroutine returned after writing the field, the write disappeared?? The strangest thing is that Emscripten's C++-to-WASM cross-compiler worked perfectly fine with the exact same routines. I wonder if the compiler I used simply had fuckass issues, it was an oldish (by modern standards) Apple clang from like probably macOS 10.14 or so.
        • VorpalWay 3 hours ago
          That or you hit some UB elsewhere that caused the compiler to assume things that you didn't follow.

          Trying UBSAN and ASAN might have been worth it. I don't think comparing debug and release builds would have helped: could be buggy optimisations or UB in your code regardless of what the outcome of that test was.

          • LoganDark 2 hours ago
            In this case I was reading memory with a debugger and could see that it literally was the same object at the same memory address, but what I did not know was whether or not different parts of the code had different views of that field for some reason, or any number of other things that could have been causing the issue. Either way, it was super weird and frustrating.
            • throw324523 1 hour ago
              The right tool for that occasion would be a watchpoint.

              1. You check in both places in code in the debugger that the address of the field matches. (Maybe the object is at the same address, but due to some build system level ifdef mismatch, you effectively get different struct definitions in different places or something like that.)

              2. You add a breakpoint inside the procedure and outside it.

              3. When the breakpoint inside is hit, you add a watchpoint at the address of the field. Continue. See if the watchpoint gets hit before reaching the second breakpoint.

    • keel_dev 29 minutes ago
      [dead]