It is straightforward to falsify theories that make clear predictions, should they happen to be wrong. It is impossible to falsify ideas that involve invisible components.

Specific theories of dark matter (e.g., WIMPs) can be seen to be increasingly unlikely, but can they be falsified outright? If we decide that WIMPs have been practically falsified – to all intents and purposes – that doesn’t mean dark matter is wrong, just that we’ve spent the past four decades building dozens of giant experiments costing billions of dollars and frustrating thousands of careers barking up the wrong tree. The unseen forest that is the ‘dark sector’ is vast; there are limitless opportunities to bark up other wrong trees.

How could we know if the entire dark sector is a non-entity? On many occasions, I’ve had colleagues say to me that they would “only consider MOND as a last resort.” OK, so when do we know we’ve reached that point? After an eternity of searching for unseen, undetected non-entities seems a little late.

This is the bind we’re in. Most of us (scientists working in the field) are unwilling to consider something as radical-seeming as MOND until dark matter has been falsified. But dark matter cannot be falsified. So we patch up our hypotheses accordingly, rinse and repeat with every new crisis#. This has been going on for nearly my entire career, to the point that I now see junior scientists who seem to think this is how science is supposed to work. Why wouldn’t they? They’ve known nothing else.

I empathize. I started from the same place: there has to be dark matter, it has to be non-baryonic, it is almost certainly a new particle and that new particle is almost certainly a WIMP. It was the hardest thing to realize and accept that we could be wrong, wrong, wrong, and wrong on all counts. I worked incredibly hard to avoid that conclusion. I take solace only in the fact that at every epoch in human history we’ve had a cosmology that we were absolutely sure$ was right but that later turned out not to be. Maybe we’re a blessed generation who finally got it right. Or maybe we’re just the latest in a long line of brilliant savants who fooled themselves into thinking we understand more than we actually do.

I remain unwilling to say that dark matter has been falsified because I don’t think it is falsifiable. That, in itself, is usually considered to be a bad thing for a scientific theory. Perhaps at some future point it will be portrayed that way for dark matter, but in the present I’ve heard plenty of scientists pretend like it is somehow a good thing. Certainly it makes a fertile playground for theorists, so I guess good/bad is a matter of perspective. I suppose an infinite forest of trees in a conveniently unobservable dark sector is an irresistible opportunity for a dog! with an insatiable bladder.

I do, however, think that dark matter has been practically falsified. If I’m wrong about that, it is possible to demonstrate. To be clear, I don’t just mean WIMPs. I mean dark matter as an explanation for the mass discrepancies observed in extragalactic systems. No kind of dark matter* can suffice. I came to this conclusion before I really knew anything about MOND, which is why I was receptive to it. Others haven’t had that experience, so aren’t.

I don’t expect other scientists to accept that an entire paradigm is wrong because I say so. I do, however, expect them to acknowledge that the dark matter hypothesis should be falsifiable. It is incumbent on each scientist to establish for themselves criteria by which that conclusion could be reached, should it be appropriate.

For me, the problem is the contradiction present in the dynamical data for galaxies. We simultaneously require galaxy disks to be maximal and also for them not to be maximal. The only way to avoid this contradiction is to engage in fine-tuning: one must build a model in which everything works out just so. Since dark matter is not otherwise falsifiable, a requirement for fine-tuning& is pretty much the worst thing we can say about it.

The contradiction as framed by others

The contradiction that concerns me has certainly been noticed by others. Usually, they choose to come down on one side or the other. The problem is that they’re both right.

On the one hand, it is clear observationally that the luminous mass matters to the dynamics of galaxies. For example, Swaters et al. note that

the luminous mass dominates the gravitational potential in the central regions, even in low surface brightness dwarf galaxies

which is to say, disk galaxies are maximal. We’ll explore what that means below, but you can see the effect by eye:

Rotation curves color coded by galaxy surface brightness. There is an almost perfect rainbow variation that is plain to see. The gravitational potential traced by the rotation curve correlates well with the distribution of stellar mass. (Adapted from McGaugh 2020.)

On the other hand, there are no residuals from the Tully-Fisher relation.

Residuals from the Tully-Fisher relation as a function of size at a given mass. Blue points are star-dominated galaxies; cyan points are gas-dominated. Compact galaxies are to the left, diffuse ones to the right. The red dashed line [δlog(V)/δlog(R) = -1/2] is what should happen with maximal disks: more compact galaxies should rotate faster at a given mass. This is what Newton predicts, but galaxies didn’t get the memo. (Adapted from McGaugh 2005.)

The inference from this observation is that galaxy disks cannot be maximal. As Courteau & Rix put it,

The case of δlog(V)/δlog(R) = -0.5 expected% for a maximal disk is ruled out

which is to say, disk galaxies are not maximal – even the high surface brightness (HSB) disks that dominated their sample.

They say this because maximal disks – or even those that contribute noticeably to the mass budget at small radii – predict deviations from the Tully-Fisher relation. They did this by intentionally focusing on the point where the stars should contribute the most. I’ve shown similar things many times, so here let’s let LCDM advocate Frank van den Bosch demonstrate it:

Model galaxies (left) of the same mass but with different spin parameters (van den Bosch 2001). In this model, spin determines size and surface brightness: lower spin galaxies are more compact and have higher surface brightness. This causes them to evince higher velocities that should cause residuals from Tully-Fisher, but these are not observed (right, McGaugh 2021).

So galaxies cannot be maximal. Only they have to be to have the observed correlation between the luminous mass distribution and the kinematics, as seen in the first figure above. Well, which is it? Are all galaxies maximal? Or none? Or can they both^ be right?

Maximal disks

Maximal disk is a technical term that is well known to people who work on the subject but not outside that relatively narrow&& field. So what do we mean by this term?

The rotation curve of NGC 6946. The solid blue line shows the expected rotation curve for the visible mass including both stars (dashed line: stellar disk; dot-dashed line: bulge) and atomic gas (dotted line). The stars suffice to describe the inner rotation curve, illustrating the case of maximum disk. The red points are the rotation curve measured optically (Daigle et al. 2006; Epinat et al. 2008); the orange points that measured with radio interferometric observations of the gas (Boomsma et al. 2008; I’ve suppressed the overlap for clarity). The thin line is the prediction of MOND; the black line is the implied dark matter distribution.

In essence, a maximal disk is one in which the stars provide practically all the mass at small radii. The depiction of NGC 6946 above illustrates maximum disk. Sure, the rotation curve flattens out at large radii and we need to invoke dark matter. But the observed stars explain the amplitude and shape of the inner rotation curve quite well. In this case there is a compact bulge at the center of the galaxy that causes a sharp rise in the rotation curve right from R = 0. (This is an example of Renzo’s Rule.) The disk (plus bulge) is maximal in the sense that we cannot attribute any more mass to them without exceeding the observed rotation curve.

The amplitude of the portion of the rotation curve due to the stars depends on their mass-to-light ratio. While this cannot exceed maximum disk, it could be lower. So conceivably, the good match to the shape of the observed rotation curve is a chimera, and really this galaxy is dark matter dominated. That is a possibility many seem to have embraced, but while it might work for the disk, it does not work for the bulge. As we suppress the contribution of the stars as Courteau & Rix argue we must, then the rotation curve looks more and more like that of the dark matter halo alone. That goes up and flattens out (and ultimately must turn over again somewhere beyond the edge of the data) but it has no features.

The rotation curve of NGC 6946 as above but with the stellar mass reduced by a factor of two. This disk is submaximal, contributing less than the implied dark matter.

One thing the rotation curve due to the dark matter halo cannot do is go up then down then up again. Yet that is exactly what it needs to do to explain the inner peak inside 1 kpc if the bulge component is not maximal. I suppose we could have a smaller dark matter halo inside the main dark matter halo that does this, but its mass distribution would have to be practically identical to that of the bulge. That’s insane. Why would we invoke a second dark matter halo when the stars are right there?

In this case, the stars have the right mass for what we expect from stellar populations. There was a long debate historically about whether the optical band mass-to-light ratios for maximum disk were consistent with those expected from stellar population synthesis models. For a long time they looked close but a bit high. This difference has pretty much gone away now that we have access to near-infrared data: the two are consistent, and having a stellar mass-to-light ratio much below the maximum disk value for high surface brightness disks becomes problematic from a population perspective.

Indeed, the 3.6 micron M*/L = 0.37 M/L in the case depicted for NGC 6946 with a maximum disk. That’s reasonable but on the low side for what is plausible for the stars in a mature spiral galaxy like this. Halving that strains credulity, so there is no room for a second inner halo, or even for the cusp predicted** for the primary cold dark matter halo. Stellar mass really does seem to dominate in the inner parts, just as Swaters et al. said.

LSB galaxies

The issue that confounded me was whether the low surface brightness (LSB) galaxies I was working on were maximal or not. My inital expectation was that LSB galaxies would be stretched out versions of HSB galaxies. I expected them to shift off of the Tully-Fisher relation and follow the line δlog(V)/δlog(R) = -0.5. They did not do that. If I didn’t have them be maximal, I found that I could explain pretty much any slope other than the one observed (δlog(V)/δlog(R) = 0). That required fine-tuning to perfectly balance the lesser contribution of stars in LSB galaxies which we had to back fill with dark matter just so. I spent ages running around in circles trying to make that work. Every time I thought I had succeeded, I realized I had assumed something that made it so: tautologies abound.

If we want to explain the shapes of rotation curves as seen up top, we need the stars to contribute to the gravitational potential. For that to work for LSB galaxies, we have to turn maximum disk up to eleven:

The low surface brightness, gas rich dwarf DDO 154. The lines have the same meaning as above, with the case of maximum disk in the left panel and that natural for stellar populations at right. The difference in disk mass is a factor of ten.

A crazy-high stellar mass-to-light ratio is what happens if we just ignore what we know about stars and just focus on the kinematics. But we do know a lot about stars. Population models indicate stellar masses that are very submaximal. Even boosting the mass-to-light ratio doesn’t get us very far. LSB galaxies aren’t really maximal in the same sense as HSB galaxies, and there is even less room for the expected cuspy halos that are already problematic when the stellar contribution is small.

Fine tuning is unavoidable

Even if we ignore what we know about stars, we still have a fine-tuning problem. The lack of a shift in the Tully-Fisher relation with either surface brightness or radial size implies that disks are all the same mass surface density. So we observe a wide range of surface brightness, but the surface mass density is always the same. That makes no sense, and is just another example of squeezing the toothpaste tube: we can make a model look OK from one perspective as long as we don’t look from another.

Worse, we still need to explain the role of the luminous mass in LSB galaxies. These are dark matter dominated at almost all radii, and yet the distribution of the observed stars and gas is predictive of the kinematics. This is a contradiction to Newtonian dynamics. The only theory that does this right – and predicted it a priori – is MOND. But that’s too horrible to contemplate, so we shield our eyes and ignore%% one or the other set of inconvenient facts. As a result, the field has become moribund, and will remain so until we free ourselves of our invisible demons.


#We’ve experienced so many crises that we seem no longer able to recognize new ones. JWST observations of high redshift galaxies follows a well-worn trajectory: an observation that contradicts the standard model is made, much huffing and puffing ensues, the theorists get to work constructing implausible models, these are accepted as patching up the hypothesis (whether satisfactory or not), and the field moves on as if nothing happened.

$To give one historical example, prior to Hubble’s discoveries in the 1920s, it was thought that the Milky Way was the entire universe. Certainly there were no other galaxies comparable to the Milky Way:

“No competent thinker, with the whole of the available evidence before him, can now, it is safe to say, maintain any single nebula to be a star system of coordinate rank with the Milky Way. A practical certainty has been attained that the entire contents, stellar and nebular, of the sphere belong to one mighty aggregation.” [i.e., the Milky Way]

-Agnes Mary Clerke in The System of the Stars (1890)

!It used to be that one would not claim a detection of dark matter until all astrophysical alternatives had been exhausted. Now it seems to be the fad to claim a detection first on the off-chance it works out later. I already peed on that tree! It’s mine!

*Excepting some sort of hybrid “dark matter” that is invented to do what ordinary dark matter cannot. By ordinary I mean CDM, WDM, SIDM, and every other variation on particle physics that simply invents new mass with no consideration of how the observed galaxy dynamics comes about. That would include primordial black holes and various macroscopic DM ideas (e.g., MACHOS, strange nuggets). Coming up with half-baked ideas for new particle dark matter is big business these days, but any idea not informed by observed astrophysics (which are most of them) is doomed to fail.

Examples of hybrid dark matter that are informed by observed astrophysics include dipolar dark matter and superfluid dark matter. Regardless of whether these specific cases are viable, the point is that the observed dynamics are a fundamental aspect if nature and require a commiserate explanation. Simply throwing in some extra mass with some fine-tuned feedback models can never provide a satisfactory explanation. Note that coming up with extra mass is mostly done by particle phenomenologists while feedback models are built by numerical astrophysicists. There is very little overlap between these communities; they pretty much just take it on faith that since dark matter has to exist, the part they don’t know about will magically work out.

&The classic example of fine-tuning in the sense that I mean is the Ptolemaic model of epicycles and deferents. If one adds enough of these and tunes them just so, anything can be fit. Note that epicycles are not explicitly falsifiable for this reason; we rejected them because they got ridiculously complicated and there turns out to be a more parsimonious explanation. The same thing holds now for dark matter and MOND.

%This slope is expected because Newton teaches us that V2 = GM/R. δlog(V)/δlog(R) = -0.5 follows from taking the logarithm of this at fixed mass. Galaxies are observed to span a large range of radius at a given mass, but not a corresponding range in circular velocity.

^Yes, they can both be right, but not with dark matter. Only MOND naturally explains both observations simultaneously.

Also, for the hyper vigilant, Courteu and Rix (1999) use a slightly different definition of velocity than I do in this Tully-Fisher residual plot. I went through all that in McGaugh & de Blok (1998) and in McGaugh (2005) and it makes no difference to the discussion here.

&&A vote we held at a conference on disk dynamics in Rome in 2000. The question of whether disks were maximal was posed; most people voted no based on the statistical lack of residuals from Tully-Fisher. After the vote, one of the dissenters noted that those who voted in favor of maximal disks were the people who actually worked in the subject. Those of us with other concerns were persuaded by the statistical evidence because we didn’t engage with the details of real, individual galaxies in the same way.

**The NFW halo famously gets the inner shape of the rotation curve wrong (the cusp-core problem), but it is also wrong at intermediate radii and at large radii. Other than that it’s great.

%%A common excuse I here for this behavior is that galaxies are “small” and nonlinear – complicated entities that we can never hope to understand, so whatever they do can be ignored as irrelevant. As a scientific argument, that’s pathetic. Galaxies should be complicated in LCDM, but in observational reality they’re kinematics are sufficiently simple that they obey a single effective force law. That’s one thing they should not do, just as a complicated set of epicycles and deferents shouldn’t always add up to the inverse square law.

10 thoughts on “The contradiction in our dark matter

    1. Frame dragging is a real effect, but it is far too weak to explain the observed mass discrepancies, especially in LSB galaxies. Besides, if I recall right, it is frequency-dependent, not acceleration-dependent as the data require. Frequency dependent hypotheses, like those that are distance-dependent, can be excluded as the first order effect https://arxiv.org/abs/astro-ph/0403610

  1. I think there is a more direct way to interpret the contradiction described here.

    General Relativity was developed and first calibrated in a very specific gravitational context: the Solar System, a system dominated by one central mass. Its structure and useful symmetries reflect this highly redundant regime. Later precision tests, including binary systems, remained comparatively low-complexity gravitational systems.

    A galaxy is not a larger Solar System. Its organization is qualitatively different. There is no dominant central mass. Instead, billions of stars form a collective, approximately axisymmetric, rotationally supported system.

    And a different phenomenology appears.

    Galaxy dynamics organize around remarkably tight relations such as the RAR and BTFR. The baryonic distribution remains strongly predictive of the observed dynamics, even where Newtonian or GR analysis says that an unseen component should dominate.

    From this perspective, dark matter can be interpreted as an ad hoc attempt to preserve GR outside the structural context where its gravitational description was empirically established. The required baryon-dark matter fine-tuning discussed in this post is then not surprising. The additional component must reproduce regularities belonging to the galaxy as an organized system.

    There are also two complementary mathematical reasons to expect such contextual limits.

    First, Algorithmic Information Theory tells us that increasing effective complexity reduces compressibility. Complexity here does not mean size or number of components. Symmetry and redundancy reduce effective complexity. When the organization changes, previously irrelevant degrees of freedom can become relevant, and the old compressed description can fail.

    Second, a physical theory is itself an information-compression scheme. It replaces many observations with a small set of equations and parameters by exploiting regularities in its domain. Change those regularities, and there is no mathematical reason to expect the same compression to remain effective.

    So the relevant transition is not simply from small systems to large systems. It is from one structural regime to another.

    GR works extraordinarily well where its characteristic structural redundancies are present. Galaxies exhibit different structural regularities, and MOND appears to capture them remarkably well.

    The relevant boundary may therefore be complexity defined by structure, not physical scale.

    A similar argument applies to the transition from galaxies to galaxy clusters.

    “Clusters ruin everything” for MOND in much the same way that galaxies ruin everything for General Relativity. In both cases, the difficulty appears when we cross into a different structural regime.

    What is being ruined is not the theory itself, but the assumption that its effective domain must be universal.

  2. Yes, dark matter is an ad hoc attempt to preserve GR. This observation alone makes no relevant predictions, so it does not follow that “The required baryon-dark matter fine-tuning discussed in this post is then not surprising” because yes, it is surprising. Just asserting that it isn’t does not a satisfactory explanation make.
    I do agree that the problem occurs beyond the realm where GR has been experimentally demonstrated to work well, so in your words GR is beyond its domain of applicability. That’s why we want a deeper theory that incorporates both GR and MOND in the appropriate limits. Whether this has anything to do with what you call “structural domains” remains to be demonstrated.

    1. A deeper theory incorporating both GR and MOND would still have to confront MOND’s shortcomings in galaxy clusters. Making MOND compatible with GR does not explain why its remarkable galaxy phenomenology becomes insufficient at the next structural transition.

      That repeated pattern is what I find significant:

      centrally dominated star systems → rotationally organized galaxies → dispersion-dominated galaxy clusters.

      At each transition, the organization, symmetries, and effective degrees of freedom change, different phenomenologies. The previous successful description then encounters problems.

      This is not an unusual way to model physics. In quantum many-body systems, effective descriptions tailored to specific organizational regimes are standard practice. Different phases and collective structures require different effective models. This is also a field with very short experimental feedback loops keeping theory in check.

      The repeated coincidence between changes in physical organization and changes in effective phenomenology is too obvious to ignore.

      1. It’s true that neither GR nor MOND provide a satisfactory explanation for rich galaxy clusters, so a joint theory might share that shortcoming. This also might be a clue to the nature of the deeper theory (e.g., eMOND). It might simply be that our baryon census remains incomplete in clusters. That everything else about clusters looks like MOND (the M-T relation parallels the BTFR, their peculiar velocities are high, they appear early in the universe) suggests that some more missing baryons is a serious possibility.
        The fact that galaxy clusters are pressure supported is not unique to them. Giant ellipticals and dwarf Spheroidals are also pressure supported systems – organized like clusters – that behave as MOND predicts. So if there is a pattern, it is not the one you point out.

        1. That is a good counterexample to my specific characterization. Pressure support alone cannot define the boundary.

          But the counterexample points to the broader structural claim. A dwarf spheroidal and a giant elliptical can both be pressure supported. They are still very different physical systems. A giant elliptical and a rich cluster are pressure supported too. They are still not structurally equivalent.

          An elliptical is one self-gravitating stellar system. A cluster is a nested system of self-gravitating galaxies, embedded in a hot intracluster medium, with several coupled dynamical components and multiple levels of organization.

          So rotation versus pressure support was too crude a variable. The relevant variable is the full organization of the system: hierarchy, acceleration regime, symmetries, redundancies, couplings between components, and effective degrees of freedom.

          This makes the cluster case more interesting, not less. Clusters retain MOND-like scaling relations (M-T tracking BTFR) while MOND becomes insufficient to describe their full dynamics. Some regularities from the galaxy regime survive into the next hierarchical level. They are no longer sufficient for a complete effective description at that level.

          That is closer to what I mean by structural domains. Scale and support type are single properties. Structure is the set of relationships between components that fixes the effective phenomenology.

          Pressure support was a weak proxy for that. What clusters likely need is a genuine M-MOND: an extension carrying its own set of parameters, fixed by cluster phenomenology, that reduce to a0 in the galaxy limit rather than borrowing a0 unchanged.

          Candidate cluster-native parameters include:

          Hot gas mass fraction. Clusters are baryon-dominated by diffuse gas, not stars. A parameter tied to the ICM mass fraction, not just total baryonic mass, would capture what actually sources the discrepancy at this level.

          A cluster-specific acceleration scale. The residual mass discrepancy in clusters is roughly constant in a way that differs from a0. If real, this points to a second characteristic acceleration set by cluster physics, not a failure of the galaxy-scale a0.

          ICM temperature and entropy profile. The M-T relation already tracks BTFR-like scaling. A parameter built from the temperature or entropy structure of the gas, rather than galaxy kinematics alone, would test whether that scaling is doing real work or is coincidental.

          Degree of hierarchical nesting. A cluster is a system of systems: galaxies within a shared potential within a gas halo. A parameter for the number of coupled dynamical components, or the depth of that hierarchy, would directly test whether hierarchy itself is the relevant structural variable.

          Dynamical state. Many clusters are unrelaxed, still merging, not virialized. A parameter for departure from equilibrium would separate genuine cluster physics from contamination by systems caught mid-transition.

          If M-MOND is right, these parameters should be derivable from cluster observables independently of galaxy data, the same way a0 was derived from disk kinematics. If the residual discrepancy in clusters turns out to correlate with gas mass fraction specifically, that would also be consistent with your missing-baryons account, without requiring the two explanations to be mutually exclusive.

          1. Maybe. It is tempting to have another relevant quantity besides acceleration, but also a narrow window to shoot for. There’s some indication of an overlap in that poorly sampled range in mass where groups sit on the BTFR but X-ray clusters deviate. So it is tempting to think it might have something to do with the presence of that ICM, but it’s harder to come up with something that works.

  3. “$To give one historical example, prior to Hubble’s discoveries in the 1920s, it was thought that the Milky Way was the entire universe. Certainly there were no other galaxies comparable to the Milky Way:”

    The Shapley-Curtis debate held on April 26, 1920, at the U.S. National Museum in Washington, D.C., between Harlow Shapley and Heber Curtis regarding the scale of the universe and the nature of spiral nebulae is the point at which the existence of other galaxies comparable with the Milky Way became a scientific theory. Both Shapley and Curtis were part right and part wrong. Shapley was right in that he placed the centre of the Milky Way at the centre of the system of globular clusters; Curtis was right in his assertion that the galaxies were like the Milky Way, not nebulae within it.

    Also Vesto Slipher started measuring velocities of galaxies at the Lowell Observatory in 1912 and by 1917 had measured 25, of which 21 were receding. So the evidence for galaxies being outside the Milky Way was available a full decade before Hubble. Indeed as the Clarke refractor telescope used by Slipher had been built in 1896 this discovery could have been made a couple of decades earlier if it had not been for Percival Lowell’s obsession with Martian canals.-

    1. Slipher’s velocities were utilized by Hubble, so maybe we’d call it Slipher’s Law had the order been reversed. The IAU recently voted to call it the Hubble-Lemaitre Law which I voted against because I thought Slipher-Hubble would have been more appropriate. They were the observers, after all. Lemaitre was the theorist who seemed to be the only one who understood what Einstein’s theory predicted for the universe, at points in the ’20s apparently including Einstein himself.
      In addition to the Great Debate – a detour I chose not to take in this post – there is a nice quote from Shapley upon receiving a letter from Hubble describing his discovery of a Cepheid in Andromeda: “Here is the letter that destroyed my universe.”

Leave a Reply

Your email address will not be published. Required fields are marked *