Full speed in reverse!

Full speed in reverse!

People have been asking me about comments in a recent video by Sabine Hossenfelder. I have not watched it, but the quote I’m asked about is “the higher the uncertainty of the data, the better MOND seems to work” with the implication that this might mean that MOND is a systematic artifact of data interpretation. I believe, because they consulted me about it, that the origin of this claim emerged from recent work by Sabine’s student Maria Khelashvili on fitting the SPARC data.

Let me address the point about data interpretation first. Fitting the SPARC data had exactly nothing to do with attracting my attention to MOND. Detailed MOND fits to these data are not particularly important in the overall scheme of these things as I’ll discuss in excruciating detail below. Indeed, these data didn’t even exist until relatively recently.

It may, at this juncture in time, surprise some readers to learn that I was once a strong advocate for cold dark matter. I was, like many of its current advocates, rather derisive of alternatives, the most prominent at the time being baryonic dark matter. What attracted my attention to MOND was that it made a priori predictions that were corroborated, quite unexpectedly, in my data for low surface brightness galaxies. These results were surprising in terms of dark matter then and to this day remain difficult to understand. After a lot of struggle to save dark matter, I realized that the best we could hope to do with dark matter was to contrive a model that reproduced after the fact what MOND had predicted a priori. That can never be satisfactory.

So – I changed my mind. I admitted that I had been wrong to be so completely sure that the solution to the missing mass problem had to be some new form of non-baryonic dark matter. It was not easy to accept this possibility. It required lengthy and tremendous effort to admit that Milgrom had got right something that the rest of us had got wrong. But he had – his predictions came true, so what was I supposed to say? That he was wrong?

Perhaps I am wrong to take MOND seriously? I would love to be able to honestly say it is wrong so I can stop having this argument over and over. I’ve stipulated the conditions whereby I would change my mind to again believe that dark matter is indeed the better option. These conditions have not been met. Few dark matter advocates have answered the challenge to stipulate what could change their minds.

People seem to have become obsessed with making fits to data. That’s great, but it is not fundamental. Making a priori predictions is fundamental, and has nothing to do with fitting data. By construction, the prediction comes before the data. Perhaps this is one way to distinguish between incremental and revolutionary science. Fitting data is incremental science that seeks the best version of an accepted paradigm. Successful predictions are the hallmark of revolutionary science that make one take notice and say, hey, maybe something entirely different is going on.

One of the predictions of MOND is that the RAR should exist. It was not expected in dark matter. As a quick review of the history, here is the RAR as it was known in 2004 and now (as of 2016):

The radial acceleration relation constructed from data available in 2004 and that from 2016.

The big improvement provided by SPARC was a uniform estimate of the stellar mass surface density of galaxies based on Spitzer near-infrared data. These are what are used to construct the x-axis: gbar is what Newton predicts for the observed mass distribution. SPARC was a vast improvement over the optical data we had previously, to the point that the intrinsic scatter is negligibly small: the observed scatter can be attributed to the various uncertainties and the expected scatter in stellar mass-to-light ratios. The latter never goes away, but did turn out to be at the low end of the range we expected. It could easily have looked worse, as it did in 2004, even if the underlying physical relation was perfect.

Negligibly small intrinsic scatter is the best one can hope to find. The issue now is the fit quality to individual galaxies (not just the group plot above). We already know MOND fits rotation curve data. The claim that appears in Dr. Hossenfelder’s video boils down to dark matter providing better fits. This would be important if it told us something about nature. It does not. All it teaches us about is the hazards of fitting data for which the errors are not well behaved.

While SPARC provides a robust estimate of gbar, gobs is based on a heterogeneous set of rotation curves drawn from a literature spanning decades. The error bars on these rotation curves have not been estimated in a uniform way, so we cannot blindly fit the data with our favorite software tool and expect that to teach us something about physical reality. I find myself having to say this to physicists over and over and over and over and over again: you cannot trust astronomical error bars to behave as Gaussian random variables the way one would like and expect in a controlled laboratory setting.

Astronomy is not conducted in a controlled laboratory. It is an observational science. We cannot put the entire universe in a box and control all the variables. We can hope to improve the data and approach this ideal, but right now we’re nowhere near it. These fitting analyses assume that we are.

Screw it. I really am sick of explaining this over and over, so I’m just going to cut & paste verbatim what I told Hossenfelder & Khelashvili by email when they asked. This is not the first time I’ve written an email like this, and I’m sure it won’t be the last.


Excruciating details: what I said to Hossenfelder & Khelashvili about the perils of rotation curve fitting on 22 September 2023 in response for their request for comments on the draft of the relevant paper:

First, the work of Desmond is a good place to look for an opinion independent of mine. 

Second, in my experience, the fit quality you find is what I’ve found before: DM halos with a constant density core consistently give the best fits in terms of chi^2, then MOND, then NFW. The success of cored DM halos happens because it is an extremely flexible fitting function: the core radius and core density can be traded off to fit any dog’s leg, and is highly degenerate with the stellar M*/L. NFW works less well because it has a less flexible shape. But both work because they have more parameters [than MOND].

Third, statistics will not save us here. I once hoped that the BIC would sort this out, but having gone down that road, I believe the BIC does not penalize models sufficiently for adding free parameters. You allude to this at the end of section 3.2. When you go from MOND (with fixed a0 it has only one parameter, M*/L, to fit to account for everything) to a dark matter halo (which has at a minimum 3 parameters: M*/L plus two to describe the halo) then you gain an enormous amount of freedom – the volume of possible parameter space grows enormously. But the BIC just says if you had 20 degrees of freedom before, now you have 22. That does not remotely represent the amount of flexibility that represents: some free parameters are more equal than others. MOND fits and DM halo fits are not the same beast; we can’t compare them this way any more than we can compare apples and snails. 

Worse, to do this right requires that the uncertainties be real random errors. They are not. SPARC provides homogeneous mass models based on near-IR observations of the stellar mass distribution. Those should be OK to the extent that near-IR light == stellar mass. That is a decent mapping, but not perfect. Consequently, we expect the occasional galaxy to misbehave. UGC 128 is a case where the MOND fit was great with optical data then became terrible with near-IR data. The absolute difference in the data are not great, but in terms of the formal chi^2 it is. So is that a failure of the model, or of the data to represent what we want it to represent?

This happens all the time in astronomy. Here, we want to know the circular velocity of a test particle in the gravitational potential predicted by the baryonic mass distribution. We never measure either of those quantities. What we measure is the (i) stellar light distribution and the (ii) Doppler velocities of gas. We assume we can map stellar light to stellar mass and Doppler velocity to orbital speed, but no mass model is perfect, nor is any patch of observed gas guaranteed to be on a purely circular orbit. These are known unknowns: uncertainties that we know are real but we cannot easily quantify. These assumptions that we have to make to do the analysis dominate over the random errors in many cases. We also assume that galaxies are in dynamical equilibrium, but 20% of spirals show gross side-to-side asymmetries, and at least 50% mild ones. So what is the circular motion in those cases? (F579-1 is a good example)

While SPARC is homogeneous in its photometry, it is extremely heterogeneous in its rotation curve measurements. We’re working on fixing that, but it’ll take a while. Consequently, as you note, some galaxies have little constraining power while others appear to have lots. That’s because many of the rotation curve velocity uncertainties are either grossly over or underestimated. To see this, plot the cumulative distribution of chi^2 for any of your models (or see the CDF published by Li et al 2018 for the RAR and Li et al 2020 for dark matter halos of many flavors. So many, I can’t recall how many CDF we published.) Anyway, for a good model, chi^2 is always close to one, so the CDF should go up sharply and reach one quickly – there shouldn’t be many cases with very low chi^2 or very high chi^2. Unfortunately, rotation curve data do not do this for any type of model. There are always way too many cases with chi^2 << 1 and also too many with chi^2 >> 1. One might conclude that all models are unacceptable – or that the error bars are Messed Up. I think the second option is the case. If so, then this sort of analysis will always have the power to mislead. 

I insert Fig. 1 from Li et al. (2020) so you don’t have to go look it up. The CDF of a statistically good model would rise sharply, being an almost vertical line at chi^2 = 1. No model of any flavor does that. That’s in large part because the uncertainties on some rotation curves are too large, while those on others are too small. The greater flexibility of dark matter models make them incrementally better than MOND for the cases with error bars that are too small – hence the corollary statement that “the higher the uncertainty of the data, the better MOND seems to work.” This happens because dark matter models are allowed to chase bogus outliers with tiny error bars in a way that MOND cannot. That doesn’t make dark matter better, it just makes it is easier to fool.

  A key thing to watch out for is the outsized effects of a few points with tiny error bars. Among galaxies with high chi^2, what often happens is that there is one point with a tiny error bar that does not agree with any of the rest of the data for any smoothly continuous rotation curve. Fitting programs penalize a model for missing this point by many sigma, so will do anything they can to make it better. So what happens is that if you let a0 vary with a flat prior, it will got to some very silly values in order to buy a tiny improvement in chi^2. Formally, that’s a better fit, so you say OK, a0 has to vary. But if you plot the fitted RCs with fixed and variable a0, you will be hard pressed to see the difference. Chi^2 is different, sure, but both will have chi^2 >> 1, so a lousy fit either way, and we haven’t really gained anything meaningful from allowing for the greater fitting freedom. Really it is just that one point that is Wrong even though it has a tiny error bar – which you can see relative to the other points, never mind the model. Dark matter halos have more flexibility from the beginning, so this is less obvious for them even though the same thing happens.

So that’s another big point – what is the prior for a dark matter halo? [Your] Table 1 allows V200 and C200 to be pretty much anything. So yes, you will find a fit from that range. For Burkert halos, there is no prior, since these do not emerge from any theory – they’re just a flexible French curve. For NFW halos, there is a prior from cosmology – see McGaugh et al (2007) among a zillion other possible references, including Li et al (2020). In any[L]CDM cosmology, the parameters V200 and C200 correlate – they are not independent. So a reasonable prior would be a Gaussian in log(C200) at a given V200 as specified by some simulation (Macio et al; see Li et al 2020). Another prior is how V200 (or M200) relates to the observed baryonic mass (or stellar mass). This one is pretty dodgy. Originally, we expected a fixed ratio between baryonic and dark mass. So when I did this kind of analysis in the ’90s, I found NFW flunked hard compared to MOND. (I didn’t know about the BIC then.) Galaxy DM halos simply do not look like NFW halos that form in LCDM and host galaxies with a few percent of their mass in the luminous disk even though this was the standard model for many years (Mo, Mao, & White 1998). If we drop the assumption that luminous galaxies are always a fixed fraction of their dark matter halos, then better fits can be obtained. I suspect your uniform prior fits have halo masses all over the place; they probably don’t correlate well with the baryonic mass, nor are their C and V200 parameters likely to correlate as they are predicted to do. You could apply the expected mass-concentration and stellar mass-halo mass relations as priors, then NFW will come off worse in your analysis because you’ve restricted them to where they ought to live.

So, as you say – it all comes down to the prior.

Even applying a stellar mass-halo mass relation from abundance matching isn’t really independent information, though that’s the best you can hope to do. But I was saying 20+ years ago that fixed mass ratios wouldn’t work, but nobody then wanted to abandon that obvious assumption. Since then, they’ve been forced to do so. But there is no good physical reason for it (feedback is the deus ex machina of all problems in the field), what happened is that the data forced us to drop the obvious assumption. Data including kinematic data (McGaugh et al 2010). So adopting a modern stellar mass-halo mass relation will give you a stronger prior than a uniform prior, but that choice has already been informed by the kinematic data that you’re trying to fit. How do we properly penalize the model for cheating about its “prior” by peaking at past data?

So, as you say – it all comes down to the prior. I think it would be important here to better constrain the priors on the DM halo fits. Li et al (2020) discuss this. Even then we’re not done, because galaxy formation modifies the form of the halo function we’re fitting. They shouldn’t end up as NFW even if they start out that way – see Li et al 2022a & b. Those papers consider the inevitable effects of adiabatic compression, but not of feedback. If feedback really has the effects on DM halos that is frequently advertised, then neither NFW or Burkert are appropriate fitting functions – they’re not what LCDM+feedback predicts. Good luck extracting a legitimate prediction from simulations, though. So we’re stuck doing what you’re trying to do: adopt some functional form to represent the DM halo, and see what fits. What you’ve done here agrees with my experience: cored DM halos work best. But they don’t represent an LCDM prediction, or any other broader theory, so – so what? 

Another detail to be wary of – the radial range over which the RC data constrain the DM halo fit is often rather limited compared to the size of the halo. To complicate matters further, the inner regions are often star-dominated, so there is not much of a handle on DM from where the data are best, at least beyond many galaxies preferring not to have a cusp since the stars already get the job done at small R. So, one ends up with V_DM(R) constrained from 3% to 10% of the virial radius, or something like that. V200 and C200 are defined at the notional virial radius, so there are many combinations of these parameters that might adequately fit the observed range while being quite different elsewhere. Even worse, NFW halos are pretty self-similar – there are combinations of (C200,V200) that are highly degenerate, so you can’t really tell the difference between them even with excellent data – the confidence contours look like bananas in C200-V200 space, with low C/high V often being as good as high C/low V. Even even even worse is that the observed V_DM(R) is often approximately a straight line. Any function looks like a straight line if you stretch it out enough. Consequently, the fits to LSB galaxies often tend to absurdly low C and high V200: NFW never looks like a straight line, but it does if you blow it up enough. So one ends up inferring that the halo masses of tiny galaxies are nearly as big as those of huge galaxies, or more so! My favorite example was NGC 3109, a tiny dwarf on the edge of the Local Group. A straight NFW fit suggests that the halo of this one little galaxy weighs more than the entire Local Group, M31 + MW + everything else combined. This is the sort of absurd result that comes from fitting the NFW halo form to a limited radial range of data. 

I don’t know that this helps you much, but you see a few of the concerns. 

Wide binary debate heats up again

Wide binary debate heats up again

One of the most interesting and contentious results concerning MOND this year has been the dynamics of wide binaries. When last I wrote on this topic, way back at the end of August, Chae (2023) and Hernandez (2023) both had new papers finding evidence for MONDian behavior in wide binaries. Since that time, they each have written additional papers on the subject. These independent efforts both report strong evidence for MONDian behavior in wide binaries, so for all of October it seemed like Game Over for conventional* dark matter.

I refrained from writing a post then because I was still waiting to see if there would be a contradictory paper. Now there is. And boy, is it contradictory! Where Hernandez et al. find 2.6σ evidence for non-Newtonian behavior and Chae finds ~5σ evidence for non-Newtonian behavior, both consistent with MOND, Banik et al. find purely Newtonian behavior and claim to exclude MOND at 19σ. That’s pretty high confidence!

Well, which is it, young feller? You got proof of non-Newtonian dynamics, or you want to insist that’s impossible?

After the latest results appeared, a red-hot debate [re]ignited on e-mail, largely along the lines of what was discussed at the conference in St. Andrews. Banik et al say that they can reproduce the MOND-like signal of Chae, but that it goes away when the data quality restriction is applied to physical velocity uncertainties (arguing that this is what you want to know) rather than to raw observational uncertainties. Chae and Hernandez counter that the method Banik et al. apply is not grounded in the Newtonian regime where everyone agrees on what should happen, so they could be calibrating the signal away. This is one thing that I had the impression that everyone had agreed to work on in St. Andrews, but it doesn’t appear that we’re there yet.

Banik et al. do a carefully planned Bayesian analysis. This approach in principle allows one to separate many effects simultaneously, one of which is close binaries (CB**). I look at the impact that close binaries have on the analysis, and it gives me the heebie-jeebies:

One panel from Fig. 10 of Banik et al.

This figure illustrates the probability of measuring a characteristic velocity in MOND for the noted range of projected sky separation. If it is just wide binaries (WB), you get the blue line. If there are some close binaries, the expected distribution changes dramatically. This change is rather larger than the signal expected from the nominal difference in gravity. You can in principle fit for everything simultaneously, but extracting the right small signal when there is a big competing signal can be tricky. Bayesian analyses can help, but they are also a double-sided sledge-hammer: a powerful tool with which to pound the data, but also a tool that can bounce back and smack you in the face. Having done such analyses, and been smacked around a few times (and having seen others get smacked around), looking at this plot really does give me the heebie-jeebies. There are lots of ways in which this can go wrong – or even just overstate the confidence of a correct result.

Everyone uses Bayesian methods these days.***

I expect people are expecting me to comment on this hot mess. Some have already asked me to do so. I really don’t want to. I’ve already said more than I should.

There are very earnest, respectable people doing this work; I don’t think anyone is being intentionally misleading. Somebody must be wrong, but it isn’t my job to sort out who. Moreover, these are long and involved analyses; it will take me time to read all the papers and make sense of them. Maybe once I do, I’ll have something more cogent to say.

I make no promises.


*By conventional dark matter, I mean new particles that only communicate with baryons via gravity.

**CB: In principle, some of the wide binaries detected by Gaia will also be close binaries, in the sense that one of the two widely separated stars is itself not a single star but an unrecognized close binary. We know this happen in nature: the nearest star system, αCentauri, is an example. The main A&B components compose a close binary with Proxima Centauri being widely separated. Modeling how often this happens in the Gaia data gives me the willies.

***To paraphrase Churchill: Many forms of statistics have been tried, and will be tried in this science of sin and woe. No one pretends+ that Bayes is perfect or all-wise. Indeed it has been said that Bayes is the worst form of statistics except for all those other forms that have been tried from time to time.

+Lots of people pretend that Bayes is perfect and all-wise.

How things go mostly right or badly wrong

How things go mostly right or badly wrong

People often ask me of how “perfect” MOND has to be. The short answer is that it agrees with galaxy data as “perfectly” as we can perceive – i.e., the scatter in the credible data is accounted for entirely by known errors and the expected scatter in stellar mass-to-light ratios. Sometimes it nevertheless looks to go badly wrong. That’s often because we need to know both the mass distribution and the kinematics perfectly. Here I’ll use the Milky Way as an example of how easily things can look bad when they aren’t.

First, an update. I had hoped to stop talking about the Milky Way after the recent series of posts. But it is in the news, and there is always more to say. A new realization of the rotation curve from the Gaia DR3 data has appeared, so let’s look at all the DR3 data together:

Gaia DR3 realizations of the Milky Way rotation curve. The most recent version of these data from Poder et al (2023) are shown as blue squares over the range 5 < R < 13 kpc. Other Gaia DR3 realizations include Ou et al. (2023, green circles), Wang et al. (2023, magenta downward pointing triangles), and Zhou et al. (2023, purple triangles).

The new Gaia realization does not go very far out, and has larger uncertainties. That doesn’t mean it is worse; it might simply be more conservative in estimating uncertainties, and not making a claim where the data don’t substantiate it. Neither does that mean the other realizations are wrong: these differences are what happens in different analyses. Indeed, all the independent realizations of the Gaia data are pretty consistent, despite the different stellar selection criteria and analysis techniques. This is especially true for R < 17 kpc where there are lots of stars informing the measurements. Even beyond that, I would say they are consistent at the level we’d expect for astronomy.

Zooming out to compare with other results:

The Milky Way rotation curve. The model line from McGaugh (2018) is shown with data from various sources. The abscissa switches from linear to logarithmic at 10 kpc to wedge it all in. The location of the Large Magellanic Cloud at 50 kpc is noted. Gaia DR3 data (Poder et al., Ou et al., Wang et al., and Zhou et al.) are shown as in the plot above. The small black squares are the Gaia DR2 realization of Eilers et al. (2019) reanalyzed to include the effect of bumps and wiggles by McGaugh (2019). Non-Gaia data include blue horizontal branch stars (light blue squares) and red giants (red squares) in the stellar halo (Bird et al. 2022), globular clusters (Watkins et al. 2019, pink triangles), VVV stars (Portail et al. 2017, dark grey squares at R < 2.2 kpc), and terminal velocities (McClure-Griffiths & Dickey 2007, 2016, light grey points from 3 < R < 8 kpc). These terminal velocities are the only data that inform the model line; everything else follows.

Overall, I would say the data paint a pretty consistent picture. The biggest tension amongst the data illustrated here is between the outermost Gaia points around R = 25 kpc and the corresponding results from halo stars. One is consistent with the model line and the other is not. We shouldn’t allow the model to inform our interpretation; the important point is that the independent data disagree with each other. This happens all the time in astronomy. Sometimes it boils down to different assumptions; sometimes it is a real discrepancy. Either way, one has to learn* to cope.

The sharp-eyed will also notice an apparent tension between the DR2 data (black squares) and DR3 around 6 and 7 kpc. This is not real – it is an artifact of different treatments of the term in the Jeans equation for the logarithmic derivative of the density profile of the tracer particles. That’s a choice made in the analysis. The data are entirely consistent when treated consistently.

Putting on an empiricist’s hat, I will say that the kink in the slope of the Gaia data around R = 18 kpc looks unnatural. That doesn’t happen in other galaxies. Rather than belabor the point further, I’ll simply say that this is how things mostly go right but also a little wrong. This is as good as we can hope for in [extra]galactic astronomy.

In contrast, it is easy to go very wrong. To give an example, here is a model of the Milky Way that was built to approximately match the rotation curve of Sofue (2020).


Fig. 1 from Dai et al. (2022). Note the logarithmic abscissa. Their caption: The rotation curve of the Milky Way. The data (solid dark circles with error bars) for r < 100kpc come from [22], while for r > 100kpc from [23]. The solid, dashed and doted lines describe the contribution from the bulge, stellar disk and dark matter halo respectively, within a ΛCDM model of the galaxy. The dashed-dot line is the total contribution of all three components.The parameters of each component are taken from [24]. For comparison, the Milky way rotation curve from Gaia DR2 is shown in color. The red dots are data from [34], the blue upward-pointing triangles are from [35], while the cyan downward-pointing triangles are from [36].

This realization of the rotation curve is very different from that seen above. Note that the rotation curve (black points) is very different from that of Gaia (red points) over the same radial range. These independent data are inconsistent; at least one of them is wrong. The data extend to very large radii, encompassing not only the LMC but also Andromeda (780 kpc away). I am already concerned about the effects of the LMC at 50 kpc; Andromeda is twice the baryonic mass of the Milky Way so anything beyond 260 kpc is more Andromeda’s territory than ours – depending on which side we’re talking about. The uncertainties are so big out there they provide no constraining power anyway.

In terms of MOND-required perfection, things fall apart for the Dai model already at very small radii. Dai et al. (2022) chose to fit their bulge component to the high amplitude terminal velocities of Sofue. That’s a reasonable thing to do, if we think the terminal velocities represent circular motion. Because of the non-circular motions that sustain the Galactic bar, they almost certainly do not – that’s why I restricted use of terminal velocities to larger radii. We also know something about the light distribution:

The inner 3 kpc of the Milky Way. The circles are the terminal velocities of Sofue (2020); the squares are the equivalent circular velocity of the potential reconstructed from the kinematics of stars in the VVV survey (Portail et al. 2017). The line is the bulge-bar model of McGaugh (2008) based on the light distribution reported by Binney et al (1997).

This is essentially the same graph as I showed before, but showing only the Newtonian bulge-bar component, and on a logarithmic abscissa for comparison with the plot of Dai et al. The two bulge models are very different. That of Dai et al. is more massive and more compact, as required to match the terminal velocities. There may be galaxies out there that look like this, but the Milky Way is not one of them.

Indeed, Newton’s prediction for the rotation curve of the bulge-bar component – the line labeled bulge/bar based on what the Milky Way looks like – is in good agreement with the effective circular speed curve obtained from stellar data. It is not consistent with the terminal velocities. We could increase the amplitude of the Newtonian prediction by increasing the mass-to-light ratio of the stars (I have adopted the value I expect for stellar populations), but the shape would still be wrong. This does not come as a surprise to most Galactic astronomers, because we know there is a bar in the center of the Milky Way and we know that bars induce non-circular motions, so we do not expect the terminal velocities to be a fair tracer of the rotation curve in this region. That’s why Portail et al. had to go to great lengths in their analysis to reconstruct the equivalent circular velocity, as did I just to build the bulge-bar model.

The thing about predicting rotation curves from the observed mass, as MOND does, is that you have to get both the kinematic data and the mass distribution right. The velocity predicted at any radius depends on the mass enclosed by that radius. So if we get the bulge badly wrong, everything spirals down the drain from there.

Dai et al. (2022) compare their model to the acceleration residuals predicted by MOND for their mass model. If all is well, the data should scatter around the constant line at zero in this graph:

Fig. 4 from Dai et al. (2022). Their caption: [The radial acceleration relation] recast as a comparison between the total acceleration, a, and the MOND prediction, aM , as a function of the acceleration due to baryons aB. The solid horizontal line is a = aM. The circles and squares with error bars represent the Milky Way and M31 data, while the gray dots are from the EAGLE simulation of ΛCDM in [1]. For aB > 10−10m/s2 any difference between a and aM is unclear. However, once aB drops well below 10−11m/s2, the discrepancy emerges. The short-dashed line is the ΛCDM fitting curve of the MW. The dash-dot line is the ΛCDM fitting curve of M31. The mass range** of galaxies in EAGLE’s data is chosen to be between 5 × 1010M to 5 × 1011M. For comparison, the Milky way rotation curve from GAIA data release II is shown in color. The red dots are data from [34], the blue triangles are from [35], while the cyan down triangles are from [36]. While the EAGLE simulation does not match the data perfectly, these plots indicate that it is much easier to accommodate a systematic downward trend with the ΛCDM model than with MOND.

Things are not well.

The interpretation that is offered (right in the figure caption) is that MOND is wrong and the LCDM-based EAGLE simulation does a better if not perfect job of explaining things. We already know that’s not right. The alternate interpretation is that this is not a valid representation of the prediction of MOND, because their mass model does not follow from the observed distribution of light. They get neither the baryonic mass distribution and its predicted acceleration ab nor the total acceleration a right in the plot above.

In terms of dark matter, the model of Dai et al. may appear viable. In terms of MOND, it is way off, not just a little off. The residuals are only zero, as they should be, for a narrow range of accelerations, 2 to 3 x 10-10 m/s/s. That’s more Newton than MOND, and appears to correspond to the limited range in radii over which their model matches the rotation curve data in their Fig. 1 (roughly 4 to 6 kpc). It doesn’t really fit the data elsewhere, and the restrictions on a MOND fit are considerably more stringent than on the sort of dark matter model they construct: there’s no reason to expect their model to behave like MOND in the first place.

And, hoo boy, does it ever not behave like MOND. Look at how far those red points – the Gaia DR2 data – deviate from zero in their Fig. 4. Those are the exact same data that agree well with the model line I show above – the data that were correctly predicted in advance. This model is a reasonable representation of the radial force predicted by MOND, with the blue line in my plot being equivalent to the zero line in theirs.

This is how things can go badly wrong. To properly apply MOND, we need to measure both the kinematics and baryonic mass distribution correctly. If we screw either up, as is easy to do in astronomy, then the result will look very wrong, even if it shouldn’t. Combine this with the eagerness many people have to dismiss MOND outright, and you wind up with lots of articles claiming that MOND is wrong – even when that’s not really the story the data tell. Happens over and over again, so the field remains stagnant.


*This is a large part of the cultural difference between physics and astronomy. Physicists are spoiled by laboratory experiments done in controlled conditions in which one can measure to the sixth place of decimals. In contrast, astronomy is an observational rather than experimental science. We can’t put the universe in a box and control all the systematics – measuring most quantities to 1% is a tall order. Consequently, astronomers are used to being wrong. While I wouldn’t say that astronomers cope with it gracefully, they’re well aware that it happens, that is has happened a lot historically, and will continue to happen in the future. It is a risk we all take in trying to understand a universe so much vaster than ourselves. This makes astronomers rather more tolerant of surprising results – results where the first response is “that can’t be right!” but also informed by the experience that “we’ve been wrong before!” Physicists coming to the field generally lack this experience and take the error bars way too seriously. I notice this attitude is creeping into the younger generation of astronomers; people who’ve received their data from distant observatories and performed CPU-intensive MCMC error analyses, so want to believe them, but often lack the experience of dozens of nights spent at the observatory sweating a thousand ill-controlled but consequential details, like walking out to a beautiful sunrise decorated by wisps of cirrus clouds. When did those arrive?!?


**The data that define the radial acceleration relation come from galaxies spanning six decades in stellar mass, so this one decade range from the simulations is tiny – it is literally comparing a factor of ten to a a factor of a million. What happens outside the illustrated mass range? Are lower masses even resolved?

Recent Developments Concerning the Gravitational Potential of the Milky Way. III. A Closer Look at the RAR Model

Recent Developments Concerning the Gravitational Potential of the Milky Way. III. A Closer Look at the RAR Model

I am primarily an extragalactic astronomer – someone who studies galaxies outside our own. Our home Galaxy is a subject in its own right. Naturally, I became curious how the Milky Way appeared in the light of the systematic behaviors we have learned from external galaxies. I first wrote a paper about it in 2008; in the process I realized that I could use the RAR to infer the distribution of stellar mass from the terminal velocities observed in interstellar gas. That’s not necessary in external galaxies, where we can measure the light distribution, but we don’t get a view of the whole Galaxy from our location within it. Still, it wasn’t my field, so it wasn’t until 2015/16 that I did the exercise in detail. Shortly after that, the folks who study the supermassive black hole at the center of the Galaxy provided a very precise constraint on the distance there. That was the one big systematic uncertainty in my own work up to that point, but I had guessed well enough, so it didn’t make a big change. Still, I updated the model to the new distance in 2018, and provided its details on my model page so anyone could use it. Then Gaia data started to pour in, which was overwhelming, but I found I really didn’t need to do any updating: the second data release indicated a declining rotation curve at exactly the rate the model predicted: -1.7 km/s/kpc. So far so good.

I call it the RAR model because it only involves the radial force. All I did was assume that the Milky Way was a typical spiral galaxy that followed the RAR, and ask what the mass distribution of the stars needed to be to match the observed terminal velocities. This is a purely empirical exercise that should work regardless of the underlying cause of the RAR, be it MOND or something else. Of course, MOND is the only theory that explicitly predicted the RAR ahead of time, but we’ve gone to great lengths to establish that the RAR is present empirically whether we know about MOND or not. If we accept that the cause of the RAR is MOND, which is the natural interpretation, then MOND over-predicts the vertical motions by a bit. That may be an important clue, either into how MOND works (it doesn’t necessarily follow the most naive assumption) or how something else might cause the observed MONDian phenomenology, or it could just be another systematic uncertainty of the sort that always plagues astronomy. Here I will focus on the RAR model, highlighting specific radial ranges where the details of the RAR model provide insight that can’t be obtained in other ways.

The RAR Milky Way model was fit to the terminal velocity data (in grey) over the radial range 3 < R < 8 kpc. Everything outside of that range is a prediction. It is not a prediction limited to that skinny blue line, as I have to extrapolate the mass distribution of the Milky Way to arbitrarily large radii. If there is a gradient in the mass-to-light ratio, or even if I guess a little wrong in the extrapolation, it’ll go off at some point. It shouldn’t be far off, as V(R) is mostly fixed by the enclosed mass. Mostly. If there is something else out there, it’ll be higher (like the cyan line including an estimate of the coronal gas in the plot that goes out to 130 kpc). If there is a bit less than the extrapolation, it’ll be lower.

The RAR model Milky Way (blue line) together with the terminal velocities to which it was fit (light grey points), VVV data in the inner 2.2 kpc (dark grey squares), and the Zhou et al. (2023) realization of the Gaia DR3 data. Also shown are the number of stars per bin from Gaia (right axis).

From 8 to 19 kpc, the Gaia data as realized by Zhao et al. fall bang on the model. They evince exactly the slowly declining rotation curve that was predicted. That’s pretty good for an extrapolation from R < 8 kpc. I’m not aware of any other model that did this well in advance of the observation. Indeed, I can’t think of a way to even make a prediction with a dark matter model. I’ve tried this – a lot – and it is as easy to come up with a model whose rotation curve is rising as one that is falling. There’s nothing in the dark matter paradigm that is predictive at this level of detail.

Beyond R > 19 kpc, the match of the model and Zhou et al. realization of the data is not perfect. It is still pretty damn good by astronomical standards, and better than the Keplerian dotted line. Cosmologists would be wetting themselves with excitement if they could come this close to predicting anything. Heck, they’re known to do that even when they’re obviously wrong*.

If the difference between the outermost data and the blue line is correct, then all it means is that we have to tweak the model to have a bit less mass than assumed in the extrapolation. I call it a tweak because it would be exactly that: a small change to an assumption I was obliged to make in order to do the calculation. I could have assumed something else, and almost did: there is discussion in the literature that the disk of the Milky Way is truncated at 20 kpc. I considered using a mass model with such a feature, but one can’t make it a sharp edge as that introduces numerical artifacts when solving the Poisson equation numerically, as this procedure depends on derivatives that blow up when they encounter sharp features. Presumably the physical truncation isn’t unphysically sharp anyway, rather being a transition to a steeper exponential decline as we sometimes see in other galaxies. However, despite indications of such an effect, there wasn’t enough data to constrain it in a way useful for my model. So rather than introduce a bunch of extra, unconstrained freedom into the model, I made a straight extrapolation from what I had all the way to infinity in the full knowledge that this had to be wrong at some level. Perhaps we’ve found that level.

That said, I’m happy with the agreement of the data with the model as is. The data become very sparse where there is even a hint of disagreement. Where there are thousands of stars per bin in the well-fit portion of the rotation curve, there are only tens per bin outside 20 kpc. When the numbers get that small, one has to start to worry that there are not enough independent samples of phase space. A sizeable fraction of those tens of stars could be part of the same stellar stream, which would bias the results to that particular unrepresentative orbit. I don’t know if that’s the case, which is the point: it is just one of the many potential systematic uncertainties that are not represented in the formal error bars. Missing those last five points by two sigma is as likely to be an indication that the error bars have been underestimated as it is to be an indication that the model is inadequate. Trying to account for this sort of thing is why the error bars of Jiao et al. are so much bigger than the formal uncertainties in the three realization papers.

That’s the outer regions. The place where the RAR model disagrees the most with the Gaia data is from 5 < R < 8 kpc, which is in the range where it was fit! So what’s going on there?

Again, the data disagree with the data. The stellar data from Gaia disagree with the terminal velocity data from interstellar gas at high significance. The RAR model was fit to the latter, so it must per force disagree with the former. It is tempting to dismiss one or the other as wrong, but do they really disagree?

Adapted from Fig. 4 of McGaugh (2019). Grey points are the first and fourth quadrant terminal velocity data to which the model (blue line) was matched. The red squares are the stellar rotation curve estimated with Gaia DR2 (DR3 is indistinguishable). The black squares are the stellar rotation curve after adjustment to be consistent with a mass profile that includes spiral arms. This adjustment for self-consistency remedies the apparent discrepancy between gas and stellar data.

In order to build the model depicted above, I chose to split the difference between the first and fourth quadrant terminal velocity data. I fit them separately in McGaugh (2016) where I made the additional point that the apparent difference between the two quadrants is what we expect from an m=2 mode – i.e., a galaxy with spiral arms. That means these velocities are not exactly circular as commonly assumed, and as I must per force assume to build the model. So I split the difference above in the full knowledge that this is not the exact circular velocity curve of the Galaxy, it’s just the best I can do at present. This is another example of the systematic uncertainties we encounter: the difference between the first and fourth quadrant is real and is telling us that the galaxy is not azimuthally symmetric – as anyone can tell by looking at any spiral galaxy, but is a detail we’d like to ignore so we can talk about disk+dark matter halo models in the convenient limit of axisymmetry.

Though not perfect – no model is – the RAR model Milky Way is a lot better than models that ignore spiral structure entirely, which is basically all of them. The standard procedure assumes an exponential disk and some form of dark matter halo. Allowance is usually made for a central bulge component, but it is relatively rare to bother to include the interstellar gas, much less consider deviations from a pure exponential disk. Having adopted the approximation of an exponential disk, one inevitably get a smooth rotation curve like the dashed line below:

Fig. 1 from McGaugh (2019). Red points are the binned fourth quadrant molecular hydrogen terminal velocities to which the model (blue line) has been fit. The dotted lines shows the corresponding Newtonian rotation curve of the baryons. The dashed line is the model of Bovy & Rix (2013) built assuming an exponential disk. The inset shows residuals of the models from the data. The exponential model does not and cannot fit these data.

The common assumption of exponential disk precludes the possibility of fitting the bumps and wiggles observed in the terminal velocities. These occur because of deviations from a pure exponential profile caused by features like spiral arms. By making this assumption, the variations in mass due to spiral arms is artificially smoothed over. They are not there by assumption, and there is no way to recover them in a dark matter fit that doesn’t know about the RAR.

Depending on what one is trying to accomplish, an exponential model may suffice. The Bovy & Rix model shown above is perfectly reasonable for what they were trying to do, which involved the vertical motions of stars, not the bumps and wiggles in the rotation curve. I would say that the result they obtain is in reasonable agreement with the rotation curve, given what they were doing and in full knowledge that we can’t expect to hit every error bar of every datum of every sort. But for the benefit of the chi-square enthusiasts who are concerned about missing a few data points at large radii, the reduced chi-squared of the Bovy & Rix model is 14.35 while that of the RAR model is 0.6. A good fit is around 1, so the RAR model is a good fit while the smooth exponential is terrible – as one can see by eye in the residual inset: the smooth exponential model gets the overall amplitude about right, but hits none of the data. That’s the starting point for every dark matter model that assumes an exponential disk; even if they do a marginally better job of fitting the alleged Keplerian downturn, they’re still a lot worse if we consider the terminal velocity data, the details of which are usually ignored.

If instead we pay attention the details of the terminal velocity data, we discover that the broad features seen there in are pretty much what we expect for the kinematic signatures of photometrically known spiral arms. That is, the mass density variations inferred by fitting the RAR correspond to spiral arms that are independently known from star counts. We’ve discussed this before.

Spiral structure in the Milky Way (left) as traced by HII regions and Giant Molecular Clouds (GMCs). These correspond to bumps in the surface density profile inferred from kinematics with the RAR (right).

If we accept that the bumps and wiggles in the terminal velocities are tracers of bumps and wiggles in the stellar mass profiles, as seen in external galaxies, then we can return to examining the apparent discrepancy between them and the stellar rotation curve from Gaia. The latter follow from an application of the Jeans equation, which helps us sort out the circular motion from the mildly eccentric orbits of many stars. It includes a term that depends on the gradient of the density profile of the stars that trace the gravitational potential. If we assume an exponential disk, then that term is easily calculated. It is slowly and smoothly varying, and has little impact on the outcome. One can explore variations of the assumed scale length of the disk, and these likewise have little impact, leading us to infer that we don’t need to worry about it. The trouble with this inference is that it is predicated on the assumption of a smooth exponential disk. We are implicitly assuming that there are no bumps and wiggles.

The bumps and wiggles are explicitly part of the RAR model. Consequently, the gradient term in the Jeans equation has a modest but important impact on the result. Applying it to the Gaia data, I get the black points:

The red squares are the Gaia DR2 data. The black squares are the same data after including in the Jeans equation the effect of variations in the tracer gradient. This term dominates the uncertainties.

The velocities of the Gaia data in the range illustrated all go up. This systematic effect reconciles the apparent discrepancy between the stellar and gas rotation curves. The red points are highly discrepant from the gray points, but the black points are not. All it took was to drop the assumption of a smooth exponential profile and calculate the density gradient numerically from the data. This difference has a more pronounced impact on rotation curve fits than any of the differences between the various realizations of the Gaia DR3 data – hence my cavalier attitude towards their error bars. Those are not the important uncertainties.

Indeed, I caution that we still don’t know what the effective circular velocity of the potential is. I’ve made my best guess by splitting the difference between the first and fourth quadrant terminal velocity data, but I’ve surely not got it perfectly right. One might view the difference between the quadrants as the level at which the perfect quantity is practically unknowable. I don’t think it is quite that bad, but I hope I have at least given the reader some flavor for some of the hidden systematic uncertainties that we struggle with in astronomy.

It gets worse! At small radii, there is good reason to be wary of the extent to which terminal velocities represent circular motion. Our Galaxy hosts a strong bar, as artistically depicted here:

Artist’s rendition of the Milky Way. Image credit: NASA/JPL-Caltech.

Bars are a rich topic in their own right. They are supported by non-circular orbits that maintain their pattern. Consequently, one does not expect gas in the region where the bar is to be on circular orbits. It is not entirely clear how long the bar in our Galaxy is, but it is at least 3 kpc – which is why I have not attempted to fit data interior to that. I do, however, have to account for the mass in that region. So I built a model based on the observed light distribution. It’s a nifty bit of math to work out the equivalent circular velocity corresponding to a triaxial bar structure, so having done it once I’ve not been keen to do it again. This fixes the shape of the rotation curve in the inner region, though the amplitude may shift up and down with the mass-to-light ratio of the stars, which dominate the gravitational potential at small radii. This deserves its own close up:

Colored points are terminal velocities from Marasco et al. (2017), from both molecular (red) and atomic (green) gas. Light gray circles are from Sofue (2020). These are plotted assuming they represent circular motions, which they do not. Dark grey squares are the equivalent circular velocity inferred from stars in the VVV survey. The black line is the Newtonian mass model for the central bar and disk, and the blue line is the corresponding RAR model as seen above.

Here is another place where the terminal velocities disagree with the stellar data. This time, it is because the terminal velocities do not trace circular motion. If we assume they do, then we get what is depicted above, and for many years, that was thought to be the Galactic rotation curve, complete with a pronounced classical bulge. Many decades later, we know the center of the Galaxy is not dominated by a bulge but rather a bar, with concominant non-circular motions – motions that have been observed in the stars and carefully used to reconstruct the equivalent circular velocity curve by Portail et al. (2017). This is exactly what we need to compare to the RAR model.

Note that 2008, when the bar model was constructed, predates 2017 (or the 2016 appearance of the preprint). While it would have been fair to tweak the model as the data improved, this did not prove necessary. The RAR model effectively predicted the inner rotation curve a priori. That’s a considerably more impressive feat than getting the outer slope right, but the model manages both sans effort.

No dark matter model can make an equivalent boast. Indeed, it is not obvious how to do this at all; usually people just make a crude assumption with some convenient approximation like the Hernquist potential and call it a day without bothering to fit the inner data. The obvious prediction for a dark matter model overshoots the inner rotation curve, as there is no room for the cusp predicted in cold dark matter halos – stars dominate the central potential. One can of course invoke feedback to fix this, but it is a post hoc kludge rather than a prediction, and one that isn’t supposed to apply in galaxies as massive as the Milky Way. Unless it needs to, of course.

So, lets’s see – the RAR model Milky Way reconciles the tension between stellar and interstellar velocity data, indicates density bumps that are in the right location to correspond to actual spiral arms, matches the effective circular velocity curve determined for stars in the Galactic bar, correctly predicted the slope of the rotation curve outside the solar circle out to at least 19 kpc, and is consistent with the bulk of the data at much larger radii. That’s a pretty successful model. Some realizations of the Gaia DR3 data are a bit lower than predicted, but others are not. Hopefully our knowledge of the outer rotation curve will continue to improve. Maybe the day will come when the data have improved to the point where the model needs to be tweaked a little bit, but it is not this day.


*To give one example, the BICEP II experiment infamously claimed in March of 2014 to have detected the Inflationary signal of primordial gravitational waves in their polarization data. They held a huge press conference to announce the result in clear anticipation of earning a Nobel prize. They did this before releasing the science paper, much less hearing back from a referee. When they did release the science paper, it was immediately obvious on inspection that they had incorrectly estimated the dust foreground. Their signal was just that – excess foreground emission. I could see that in a quick glance at the relevant figure as soon as the paper was made available. Literally – I picked it up, scanned through it, saw the relevant figure, and could immediately spot where they had gone wrong. And yet this huge group of scientists all signed their name to the submitted paper and hyped it as the cosmic “discovery of the century”. Pfft.

Recent Developments Concerning the Gravitational Potential of the Milky Way. II. A Closer Look at the Data

Recent Developments Concerning the Gravitational Potential of the Milky Way. II. A Closer Look at the Data

Continuing from last time, let’s compare recent rotation curve determinations from Gaia DR3:

Fig. 1 from Jiao et al. comparing three different realizations of the Galactic rotation curve from Gaia DR3. The vertical lines* mark the range of the Ou et al. data considered by Chan & Chung Law (2023).

These are different analyses of the same dataset. The Gaia data release is immense, with billions of stars. There are gazillions of ways to parse these data. So it is reasonable to have multiple realizations, and we shouldn’t expect them to necessarily agree perfectly: do we look exclusively at K giants? A stars? Only stars with proper motion and/or parallax data more accurate than some limit? etc. Of course we want to understand any differences, but that’s not going to happen here.

My first observation is that the various analyses are broadly consistent. They all show a steady decline over a large range of radii. Nothing shocking there; it is fairly typical for bright, compact galaxies like the Milky Way to have somewhat declining rotation curves. The issue here, of course, is how much, and what does it mean?

Looking more closely, not all of the data agree with each other, or even with themselves. There are offsets between the three at radii around the sun (we live just outside R = 8 kpc) where you’d naively think they would agree the best. They’re very consistent from 13 < R < 17 kpc, then they start to diverge a little. The Ou data have a curious uptick right around R = 17 kpc, which I wouldn’t put much stock in; weird kinks like that sometimes happen in astronomical data. But it can’t be consistent with a continuous mass distribution, and will come up again for other reasons.

As an astronomer, I’m happy with the level of agreement I see here. It is not perfect, in the sense that there are some points from one data set whose error bars do not overlap with those of other data sets in places. That’s normal in astronomy, and one of the reasons that we can never entirely trust the stated uncertainties. Jiao et al. make a thorough and yet still incomplete assessment of the systematic uncertainties, winding up with larger error bars on the Wang et al. realization of the data.

For example, one – just one of the issues we have to contend with – is the distance to each star in the sample. Distances to individual objects are hard, and subject to systematic uncertainties. The reason to choose A stars or K giants is because you think you know their luminosity, so can estimate their distance. That works, but aren’t necessarily consistent (let alone correct) among the different groups. That by itself could be the source of the modest difference we see between data sets.

Chan & Chung Law use the Ou et al. realization of the data to make some strong claims. One is that the gradient of the rotation curve is -5 km/s/kpc, and this excludes MOND at high confidence. Here is their plot.

You will notice that, as they say, these are the data of Ou et al, being identical to the same points in the plot from Jiao et al. above – provided you only look in the range between the lines, 17 < R < 23 kpc. This is where the kink at R = 17 kpc comes in. They appear to have truncated the data right where it needs to be truncated to ignore the point with a noticeably lower velocity, which would surely affect the determination of the slope and reduce its confidence level. They also exclude the point with a really big error bar that nominally is within their radial range. That’s OK, as it has little significance: it’s large error bar means it contributes little to the constraint. That is not the case for the datum just inside of R = 17 kpc, or the rest of the data at smaller radii for that matter. These have a manifestly shallower slope. Looking at the line boundaries added to Jiao’s plot, it appears that they selected the range of the data with the steepest gradient. This is called cherry-picking.

It is a strange form of cherry-picking, as there is no physical reason to expect a linear fit to be appropriate. A Keplerian downturn has velocity decline as the inverse square root of radius (see the dotted line above.) These data, over this limited range, may be consistent with a Keplerian downturn, but certainly do not establish that it is required.

Contrast the statements of Chan & Chung Law with the more measured statement from the paper where the data analysis is actually performed:

… a low mass for the Galaxy is driven by the functional forms tested, given that it probes beyond our measurements. It is found to be in tension with mass measurements from globular clusters, dwarf satellites, and streams.

Ou et al. (2023)

What this means is that the data do not go far enough out to measure the total mass. The low mass that is inferred from the data is a result of fitting some specific choice of halo form to it. They note that the result disagrees with other data, as I discussed last time.

Rather than cherry pick the data, we should look at all of it. Let’s see, I’ve done that before. We looked at the Wang et al. (2023) data via Jiao et al. previously, and just discussed the Ou et al. data. That leaves the new Zhao et al. data, so let’s look at those:

Milky Way rotation curve with RAR model (blue line from 2018) and the Gaia DR3 data as realized by Zhou et al. (2023: purple triangles). The dashed line shows the number of stars (right axis) informing each datum.

These data were the last of the current crop that I looked at. They look… pretty good in comparison with the pre-existing RAR model. Not exactly the falsification I had been led to expect.

So – the three different realizations of the Gaia DR3 data are largely consistent, yet one is being portrayed as a falsification of MOND while another is in good agreement with its prediction.

This is why you have to take astronomical error bars with a grain of salt. Three different groups are using data from the same source to obtain very nearly the same result. It isn’t quite the same result, as some of the data disagree at the formal limits of their uncertainty. No big deal – that’s what happens in astronomy. The number of stars per bin helps illustrate one reason why: we go from thousands of stars per bin near the sun to tens of stars in wider bins at R > 20 kpc. That’s not necessarily problematic, but it is emblematic of what we’re dealing with: great gobs of data up close, but only scarce scratches of it far away where systematic effects are more pernicious.

In the meantime, one realization of these data are being portrayed as a death knell for a theory that successfully predicts another realization of the same data. Well, which is it?


*Thanks to Moti Milgrom for pointing out the restricted range of radii considered by Chan & Chung Law and adding the vertical lines to this figure.

Recent Developments Concerning the Gravitational Potential of the Milky Way. I.

Recent Developments Concerning the Gravitational Potential of the Milky Way. I.

Recent results from the third data release (DR3) from Gaia has led to a flurry of papers. Some are good, some are great, some are neither of those. It is apparent from the comments last time that while I’ve kept my pledge to never dumb it down, I have perhaps been assuming more background knowledge on the part of readers than is adequate. I can’t cram a graduate education in astronomy into one web page, but will try to provide a little relevant context.

Galactic Astronomy is an ancient field, dating back at least to the Herschels. There is a lot that is known in the field. There have also been a lot of misleading observations, going back just as far to the Herschel’s map of the Milky Way, which was severely limited by extinction from interstellar dust. That’s easy to say now, but Herschel’s map was the standard for over a century – longer than our modern map has persisted.

So a lot has changed, including a lot that seemed certain, so I try to keep an open mind. The astronomers working with the Gaia data – the ones deriving the rotation curve – are simply following where those data take them, as they should. There are others using their analyses to less credible ends. A lot of context is required to distinguish the two.

The total mass of the Milky Way

There are a lot of constraints on the mass of the Milky Way that predate Gaia; it’s not like these are the first data that address the issue. Indeed, there are lots and lots and lots of other applicable data acquired using different methods over the course of many decades. Here is a summary plot of determinations of the mass of the Milky Way compiled by Wang et al. (2019).

This is an admirable compilation, and yet no such compilation can be complete. There are just so many determinations by lots of independent authors. Still, this is nice for listing multiple results from many distinct methodologies. They all consistently give numbers around 1012 solar masses. (Cast in these terms, my own estimate is 1.4 x 1012 albeit with a substantial systematic uncertainty.) I’ve added a point for the total mass according to the alleged Keplerian downturn seen in the Gaia data, 2 x 1011 solar masses. One of these things is not like the others.

The difference from the bulk of the data has nearly every astronomer rolling our collective eyes. Most of us straight up don’t believe it. That’s not to say the Gaia data are wrong, but the interpretation of those data as indicative of such a small, finite total mass seems unlikely in the light of all other results.

As I discussed briefly last time, it is conceivable that previous results are wrong or misleading due to some systematic effect or bad assumption. For example, mass estimates based on “satellite phenomenon” require the assumption that the satellite galaxies are indeed satellites of the Milky Way on bound orbits. That seems like a really good assumption, as without it, their presence is an instantaneous coincidence particular to the most recent few percent of a Hubble time: they wouldn’t have been nearby more than a billion years ago, and won’t be around another for even a few hundred million more. That sounds like a long time to you and me, but it is not that long on a cosmic scale. Maybe they’re raining down all the time to give the appearance of a steady state? Where have I heard that before?

Even if we’re willing to dismiss satellite constraints, that doesn’t suffice. It isn’t good enough to find flaw with one set of determinations; one must question all distinct methods. I could probably do that; there’s always a systematic uncertainty that might be bigger than expected or an assumption that could go badly wrong. But it is asking a lot for all of them to conspire to be wrong at the same time by the same amount. (The assumption of Newtonian gravity is a catch-all.)

Some constraints are more difficult to dodge than others. For example, the escape velocity method merely notes that there are fast moving stars in the solar neighborhood. Those stars are many billions of years old, and wouldn’t be here if the gravitational potential couldn’t contain them. The mass implied by the Gaia quasi-Keplerian downturn doesn’t suffice.

That said, the total mass of the Milky Way as expressed above is a rather notional quantity. M200 occurs roughly 200 kpc out for the Milky Way, give or take a lot. And the “200” in the subscript has nothing to do with that radius being 200 kpc for reasons too technical and silly to delve into. So my biggest concern about the compilation above is not that the data are wrong so much as they are being extrapolated to an idealized radius that we don’t directly observe. This extrapolation is usually done by assuming the potential of an NFW halo, which makes perfect sense in terms of LCDM but none whatsoever empirically, since NFW predicts the wrong density profile at small, intermediate, and large radii: where the density profile ρ ∝ r is predicted to have α = (1,2,3), it is persistently observed to be more like (0,1,2). While the latter profile is empirically more realistic, it also fails to converge to a finite total mass, rendering the concept meaningless.

Rather than indulge yet again in a discussion of the virtues and vices of different dark matter halo profiles, let’s look at an observationally more robust quantity: the enclosed mass. Wang et al. also provide a tabulation of this quantity from many sources, as depicted here:

Rotation curve constraints implied by the enclosed mass measurements tabulated by Wang et al. (2019) combined with the halo stars and globular clusters previously discussed. The location of the Large Magellanic Cloud is also indicated; data beyond this radius (and perhaps even within it) are subject to perturbation by the passage of the LMC. The RAR-based model is shown as the blue line; the light blue line includes a very uncertain estimate of the effect of the coronal gas. This is very diffuse and extended, and only becomes significant at very large radii. The dotted line is the Keplerian curve for a mass of 2 x 1011 M.

Not all of the enclosed mass data are consistent with one another. The bulk of them are consistent with the RAR model Milky Way (blue line). None of them are consistent with the small mass indicated by recent Gaia analyses (dotted line). Hence the collective unwillingness of most astronomers to accept the low-mass interpretation.

An important thing to note when considering data at large radii, especially those beyond 50 kpc, is that 50 kpc is the current Galactocentric radius of the Large Magellanic Cloud. The LMC brings with it its own dark matter halo, which perturbs the outer regions of the Milky Way. This effect is surprisingly strong*, and leads to the inference that the mass ratio of the two is only 4 or 5:1 even though the luminosity ratio is more like 20:1. This makes the interpretation of the data beyond 50 kpc problematic. If we use that as a pretext to ignore it, then we infer that our low mass Milky Way is no more massive then the LMC – an apparently absurd situation.

There are many rabbit holes we could dig down here, but the basic message is that a small Milky Way mass violates a gazillion well-established constraints. That doesn’t mean the Gaia data are wrong, but it does call into question their interpretation. So next time we’ll look more closely at the data.


*This is not surprising in MOND. The LMC is in the right place at the right time to cause the Galactic warp. The LMC as a candidate perturber to excite the Galactic warp was recognized early, but the conventional mass was thought to be much too small to do the job. The small baryonic mass of the LMC in MOND is not a problem as the long range nature of the force law makes tidal effects more pronounced: it works out about right.

Wide Binary Results Favoring MOND

I think the time has come for another update on wide binaries. These were intensely debated at the conference in St. Andrews, with opposing camps saying they did or did not show MONDian behavior. Two papers by independent authors have recently been refereed and published: Chae (2023) in the Astrophysical Journal and Hernandez (2023) in Monthly Notices. These papers both find evidence for MONDian behavior in wide binaries.

If these new results are correct, they are the smoking gun for MOND. I’ve been trying to avoid that phrase, and think of how we would explain this with dark matter. I haven’t come up with any good ideas. This doesn’t preclude others from coming up with bad ideas, but the problem this result poses is profound.

The basic idea is that galaxies reside in dark matter halos. These are diffuse entities with a particular mass distribution that must contribute the right gravitational force to explain observations on galactic scales. On local scales, like the solar neighborhood, this leads to a very low space density of about 0.007 solar masses per cubic parsec, or 0.26 GeV/cm3. For comparison, the local density of stars and gas is about 0.11 solar masses per cubic parsec. Adding up all the dark matter in the solar system within the orbit of Pluto amounts to the equivalent mass of a one km-size asteroid. That doesn’t do anything noticeable to solar system dynamics, especially when it is spread out as expected rather than concentrated in an asteroid.

Wide binaries should encompass more dark matter than the solar system by virtue of their greater size, but the enclosed mass remains too tiny to affect the orbits of the stars. There could be the occasional lump of dark matter, but those should be few and far between: the conventional expectation for binary stars is purely Newtonian, with no hint of a mass discrepancy. In contrast, the expectation in MOND is that every system that experiences the low acceleration regime should show a discrepancy of predictable amplitude. I simply don’t see how to imitate that with any of the usual dark matter suspects.

Here is the results from Chae’s paper. There are many figures like this that explore all sorts of permutations on sample selection and other effects. The answer persistently comes up the same. There is a systematic deviation from Newtonian behavior that is consistent with MOND, and in particular with the nonlinear theory AQUAL proposed early on by Bekenstein & Milgrom.

Part of Fig. 19 from Chae (2023). As one goes to lower acceleration, the data for wide binaries agrees well with the prediction of the Aquadratic Lagrangian theory of MOND (purple line in lower panel).

This figure subsumes many astronomical details, like the distribution of orbital eccentricities and the frequency of triple systems. Chae has simulated what to expect as a result of all these effects, with the results in the top panel distinguishing between the Newtonian expectation in blue and the data in red. At high accelerations, the red histogram is right on top of the blue histogram. These distributions are indistinguishable, as they should be in both theories. As one looks to lower accelerations, the red and blue histograms begin to part. They stand clearly apart in the lowest acceleration bin. This is as expected in MOND. In contrast, the histograms should never diverge in the Newtonian case, with or without dark matter.

A similar result has been obtained by Hernandez (2023), who emphasizes the importance of obtaining a clean sample for which one is sure that the binaries are genuinely bound and have radial velocities as well as proper motions. The data follow the Newtonian line until they don’t. The deviation is consistent with MOND.

Part of Fig. A1 from Hernandez (2023). The MOND effect is apparent as the break of the red points from the purely Newtonian blue line.

Again, there are many figures like this in the paper to explore all the possible permutations. These all paint the same picture: MOND. The published result Hernandez obtains is consistent with the result obtained by Chae, relieving a small tension that was present in the preprint stage.

Still outstanding is why Chae and Hernandez get a different answer from Pittordis & Sutherland (2023), who utilize many more binaries. This is a tradeoff that frequently arises in astronomical data analysis: numbers vs. quality. The risk with numbers is that the signal you’re searching for gets drowned out in a sea of noise. The risk in defining a high quality sample is that you unintentionally introduce a selection effect that causes a signal to appear where there isn’t one. It seems unlikely that this would result in MOND-like behavior – it could do any number of crazy things – but I don’t know enough about this specific subject to judge. Note that I’m willing to say when I’m out of my expertise; I expect it won’t be hard to find faux experts who don’t acknowledge the limitations of their qualifications and are perfectly happy to find flaws with studies they dislike but don’t understand.

What I hope to see in future is some convergence between the different groups, or at least for some understanding to emerge as to why their results differ. In the meantime, I expect most of the community will duck and cover.

Is NGC 1277 a problem for MOND?

Is NGC 1277 a problem for MOND?

Alert reader Dan Baeckström recently asked about NGC 1277, as apparently some people have been making this out to be some sort of death knell for MOND.

My first reaction was NGC who? There are lots of galaxies in the New General Catalog (new in 1888, even then drawing heavily on earlier work by the Herschels). I’m well acquainted with many individual galaxies, and can recall many dozens by name, but I do not know every single thing in the NGC. So I looked it up.

NGC 1277 in the Perseus cluster. Photo credit: NASA, ESA, M. Beasley, & P. Kehusmaa

NGC 1277 is a lenticular galaxy. Early type. Lots of old stars. These types of galaxies tend to be baryon dominated in their centers. One might even describe them as having a dearth of dark matter. This is expected in MOND, as the stars are sufficiently concentrated that these objects are in the high acceleration regime near their centers. The modification only appears when the acceleration drops below a0 = 1.2 x 10-10 m/s/s; when accelerations are above this scale, everything is Newtonian – no modification, no need for dark matter.

So, is NGC 1277 special in some way? Why does this come up now?

There is a recent paper on NGC 1277 by Comerón et al. that seems to be the source of the claims of a death knell. The title is The massive relic galaxy NGC 1277 is dark matter deficient. That sounds normal for this type of galaxy, but I guess if you disliked MOND without understanding it, you might misinterpret that title to mean there was no mass discrepancy at all, hence a problem for MOND. I guess. I’m an expert on the subject; I don’t know where non-experts get their delusions.

The science paper by Comerón et al. is a nice analysis of reasonably high quality observations of the kinematics of this galaxy. Not seeing what the worry is. Here is their Fig. 19, which summarizes the enclosed mass distribution:

Three-dimensional cumulative mass profiles of NGC 1277 (Fig. 19 of Comerón et al.) Stars and the central black hole account for everything within the observed radius; dark matter (colored bands) is not yet needed.

The first thing I did was eyeball this plot and calculate the circular speed of a test particle at 10 kpc near the edge of the plot. Newton taught us that V2 = GM/R, and the enclosed mass there looks to be just shy of 2 x 1011 solar masses, so V = 290 km/s. That’s big, but also normal for a massive galaxy like this. The corresponding centripetal acceleration V2/R is about 2a0. As expected, this galaxy is in the high acceleration regime, so MOND predicts Newtonian behavior. That means the stars suffice to explain the dynamics; no need for dark matter over this range of radii.

The second thing I did was check to see what Comerón et al. said about it themselves. They specifically address the issue, saying

One might be tempted to use the fact that NGC 1277 lacks detectable dark matter to speculate about the (in)existence of Milgromian dynamics (also known as MOND; Milgrom 1983) or other alternatives to the ΛCDM paradigm. Given a centrally concentrated baryonic mass of M ≈ 1.6 × 1011M and an acceleration constant a0 = 1.24 × 10−10 m s−2 (McGaugh 2011), a radius R = 13 kpc should be explored to be able to probe the fully Milgromian regime. This is about twice the radius that we cover and therefore our data do not permit studying the Milgromian regime 

Comerón et al. (2023)

which is what I just said. These observations do not probe the MOND regime, and do not test theory. So, in order to think this work poses a problem for MOND, you have to (i) not understand MOND and (ii) not bother to read the paper.

I wish I could say this was unusual. Unfortunately, it is only a bit sub-par for the course. A lot of people seem to hate MOND. I sympathize with that; I was really angry the first time it came up in my data. But I got over it: anger is not conducive to a rational assessment of the evidence. A lot of people seem to let their knee-jerk dislike of the idea completely override their sense of objectivity. All too often, they don’t even bother to do minimal fact checking.

As Romanowsky et al. pointed out, the dearth of dark matter near the centers of early type galaxies is something of a problem for the dark matter paradigm. As always, this depends on what dark matter actually predicts. The most obvious expectation is that galaxies form in cuspy dark matter halos with a high concentration of dark matter towards the center. The infall of baryons acts to further concentrate the central dark matter. So the nominal expectation is that there should be plenty of dark matter near the centers of galaxies rather than none at all. That’s not what we see here, so nominally NGC 1277 presents more of a challenge for the dark matter paradigm than it does for MOND. It makes no sense to call foul on one theory without bothering to check if the other fares better. But we seem to be well past sense and well into hypocrisy.

The MOND at 40 conference

I’m back from the meeting in St. Andrews, and am mostly recovered from the jet lag and the hiking (it was hot and sunny, we did not pack for that!) and the driving on single-track roads like Mr. Toad. The A835 north from Ullapool provides some spectacular mountain views, but the A837 through Rosehall is more perilous carnival attraction than well-planned means of conveyance.

As expected, the most contentious issue was that of wide binaries. The divide was stark: there were two talks finding nary a hint of MONDian signal, just old Newton, and two talks claiming a clear MONDian signal. Nothing was resolved in the sense of one side convincing the other it was right, but there was progress in terms of [mostly] amicable discussion, with some sensible suggestions for how to proceed. One suggestion was that a neutral party should provide all the groups with several sets of mock data, one Newtonian, one MONDian, and one something else, to see if they all recovered the right answers. That’s a good test in principle, but it is a hassle to do in practice, as it is highly nontrivial to produce realistic mock Gaia data, so no one was leaping at the opportunity to stick their hand in this particular bear trap.

Xavier Hernandez made the excellent point that one should check that one’s method recovers Newtonian behavior for close binaries before making any claims to require/exclude such behavior for wide binaries. Neither MOND nor dark matter predicts any deviation from Newtonian behavior where stars are orbiting each other well in excess of a0, of which there are copious examples, so they provide a touchstone on which all should agree. He also convinced me that it was a Good Idea to have radial velocities as well as proper motions. This limits the sample size, but it helps immensely to insure that sample binaries are indeed bound pairs of binary stars. Doing this, he finds MOND-like behavior.

Previously, I linked to a talk by Indranil Banik, who found Newtonian behavior. This led to an exchange with Kyu-Hyun Chae, who has now posted an update to his own analysis in which he finds MONDian behavior. It is a clear signal, and if correct, could be the smoking gun for MOND. It wouldn’t be the first one; that honor probably goes to NGC 1560, and there have been plenty of other smoking guns since then. The trick seems to be finding something than cannot be explained with dark matter, and this could play that role since dark matter shouldn’t be relevant to binary stars. But dark matter is pretty much the ultimate Rube Goldberg machine of science, so we’ll see explanation people come up with, should they need to do so.

At present, the facts of the matter are still in dispute, so that’s the first thing to get straight.


Thanks to everyone I met at the conference who told me how useful this blog is. That’s good to know. Communication is inefficient at best, counterproductive at worst, and most often practically nonexistent. So it is good to hear that this does some small good.

Commentary on Wide Binaries

Last time, I commented on the developing situation with binary stars as a test of MOND. I neglected to enable comments for that post, so have done so now.

Indranil Banik has shared his perspective on wide binaries in a talk on the subject that is available on Youtube, included below.

Indranil and his collaborators are not seeing a MOND effect in wide binaries. Others have, as I discussed in the previous post. After the video posted above, Indranil comments on the work of Kyu-Hyun Chae:

Regarding the article by Chae (https://arxiv.org/abs/2305.04613), equation 7 of MNRAS 506, 2269–2295 (2021) shows that the relative velocity is limited such that the v_tilde parameter (ratio of relative velocity within the sky plane to the Newtonian circular velocity at the projected separation) is at most 1 for 5 M_Sun binaries and in general is sqrt(5 M_Sun/M) for a binary of total mass M. This means v_tilde only goes up to 2 for M = 1.25 M_Sun, but more generally it goes up to a higher value at lower mass. Since the main signal in MOND is a broader v_tilde distribution at lower acceleration and a lower mass reduces the acceleration, this can lead to an artificial signal whereby lower mass systems have a larger rms v_tilde. Now a simple rms statistic is not exactly what Chae did, but this does highlight the kind of problem that can arise. Indeed, the v_tilde distribution prepared by Chae for the article in its figure 25 does show a rather sharp decline in the v_tilde distribution – there is not much of an extended tail, even less than in the model! This is obviously not due to measurement errors and contaminating effects like chance alignments, which would broaden the tail further. Rather, it is due to the upper limit to v_tilde imposed from the sample selection. This just means the underlying sample used is not well suited to the wide binary test, since it was quite clear a priori that the main signal for MOND would be in the region of v_tilde = 1-1.5 or so. One possibility is to try and restrict the analysis to a narrower range of binary total mass to try and alleviate the above concern, in which case the upper limit to v_tilde would be perhaps above 2 for the full sample used. There is however another issue in that lower accelerations generally correspond to higher separations and thus lower orbital velocities, so the fractional uncertainty in the velocity is likely to be larger. Thus, the v_tilde distribution is likely to be broader at low accelerations. This can be counteracted by having low errors across the board, but then the key quantity is the uncertainty on v_tilde. This aspect is not handled very rigorously – it is assumed that if the proper motions are accurate to better than 1%, then v_tilde will be sufficiently well known. But if the tangential velocity is about 20 km/s, a 1% error means an error of 200 m/s on the velocity of each star, so the relative velocity has an uncertainty of about 280 m/s. This is quite large compared to typical wide binary relative velocities, which are generally a few hundred m/s. Without doing a more detailed analysis, perhaps one thing to do would be to change this 1% requirement to 0.5% or 1.5% and see what happens. I am therefore not convinced that the MOND signal claimed by Chae is genuine.

I. Banik

Kyu-Hyun Chae responded to that, but apparently many people are not able to see his response on Youtube. I cannot. So I asked him about it, and he shares it here:

Since Indranil sent this concern to me in person, I’m replying here. No cut on v_tilde is used in my analysis because it is a gravity test. I did not use equation 7 of El-Badry et al. (MNRAS 506, 2269–2295 (2021)) to cut out high v_tilde data, so there are some (though relatively small number of) data points above equation (7). I removed chance alignment cases by requiring R < 0.01 (El-Badry et al. convincingly show that R can be used to effectively remove chance alignment cases). This is the main reason why there is no high velocity tail. I have already considered varying proper motion (PM) relative errors: there are three cases PM rel error < 0.01 (nominal case), <0.003 (smaller case), and <0.2 (larger case). The conclusion on gravity anomaly (MOND signal) is the same in all three cases although the fitted f_multi (multiplicity fraction) varies. We can have more discussion in the st Andrews June meeting. I’m sure it will take some time but you will be convinced that my results are correct.

K.-H. Chae

He also shares this figure:

This is how the science sausage is made. As yet, there is no consensus.