Finished! Looks like this project is out of data at the moment!

See Results

Thank you everybody for the great effort! We are sharing with you the results of your amazing work here

Results

Hello, dear detectives,

Thanks to all of you, we were able to collect over 1 million classifications in just 2.5 months. We truly appreciate your tremendous effort and active participation on the forum. It has been a delight to see so many fascinating objects, and we hope you share that excitement.

Thanks to your efforts, we have improved the detection of real astrophysical sources by 5%, and we are now able to filter out 29% more bogus sources. This new model, incorporating your labels, has been in use since the beginning of July 2026 for Rubin Observatory alerts. With Rubin expected to produce ~7 million alerts each night, your collective effort will remove ~3 million bogus events each and every night. This reduces the risk of wasting precious follow-up resources (observations with other observatories) and helps astronomers focus our efforts on real astrophysical transients.

Our primary goal was to integrate real Rubin data into the training of a machine learning model that had previously only learned from simulated data. While simulated data perform well when dealing with perfect sources on the difference image, they struggle when sources deviate from that ideal shape. Such deviations are quite common and can occur for simple reasons, many of which are beyond an astronomer's control. For example, data collected on a humid night can introduce numerous deviations!

The Rubin sources we provided were selected in various ways: those with positional matches to star catalogs (labelled “gaia”), asteroids, comets, or other moving objects in our Solar System (“ss”), as well as possible cosmic explosions (“tns”, “galaxy”), and some sources that had already been flagged as Real by the existing machine learning model. Among all the sources you classified as Real, about 40% were associated with Solar System objects, while roughly 50% of your Bogus classifications were associated with star sources, largely due to artifacts in the difference image. Overall, your classifications aligned well with our expectations from this data in particular.

Each source received between 2 and 19 classifications from different volunteers. By using the majority vote to determine the final label (Real or Bogus), we found that, in most cases, participants agreed on the same label. For Solar System objects, for example, about 70% (0.7 probability) of volunteers converged on the same classification. In contrast, sources from the Gaia catalog showed the least agreement, with only about 10% of users concurring on a label. (Spoiler alert: astronomers also tend to disagree frequently on star sources due to artifacts in the difference image.)

Your contribution to the Rubin Observatory

We followed a detailed analysis by comparing some of your classifications to those made by astronomers in the field (more detail and technical aspects of the method we followed can be seen in the scientific article: SPACE WARPS – I. Crowdsourcing the discovery of gravitational lenses). The results of this analysis helped define the final labels used to train the machine learning binary classifier, which determines whether a source is Real or Bogus.

Thanks to your efforts, we have improved the detection of real astrophysical sources by 5%, from 90% with the old model to 95% with the new one. Additionally, we are now able to filter out 29% more bogus sources, a significant leap from 62% to 91%. You help us decrease the number of False Bogus and False Real.

This new model has been in use since the beginning of July 2026 for Rubin Observatory alerts. More information and resources about these alerts can be found at https://www.zooniverse.org/projects/ebellm/rubin-difference-detectives/about/research. The observatory does not simply use the binary Real or Bogus label; instead, it provides a probability score (from 0% to 100%) indicating the probability that a source is genuinely astrophysical.

we want to have the Real-Pred Real and Bogus-Pred Bogus as close to 100% as possible by reducing the Real-Pred Bogus and Bogus-Pred Real

Thank you once again for your invaluable contribution!

Real sources we were able to recover thanks to you!

Bogus we are now able to identify!

Rubin data is changing every year, and the model needs to be fine-tuned to accommodate those changes; we are planning to do yearly rounds of this project. We hope to see you all here!