Showing posts with label optimization. Show all posts
Showing posts with label optimization. Show all posts

Thursday, January 28, 2010

Working Data Reconstruction!

I ran the SDSS CAS data though our working pipeline. I've run on a small data set as a first go -- the same size as what I was running on for the reconstructions with the mock data -- 20,000 points in both the photometric and spectroscopic data sets. The difference here is that both data sets are coming from the same super-set of data (I am randomly picking points from a larger set of Stripe-82 data), so the redshift distributions of the two sets are the same (see below plot). This, in theory, should make the reconstruction easier to do.

Redshift Distributions of photometric and spectroscopic data sets


The Bad News
The correlation functions look really noisy. This is hopefully due to the fact that I am running with 20,000 points instead of 1,000,000 points (before) so hopefully the noise will clear up when I re-run with a larger data set. I wouldn't expect the correlations to look as smooth on a smaller data set as it does with the mock data, because the mock data, by design, traces dark matter populations explicitly, so the dark matter signal is put in by hand. There is also more noise on this data than on the mock data. It is logical to me that we might need higher statistics to get the same strength of signal as we did in the mock data.


The Good News
Even with the noisy correlation functions, the reconstruction looks not too bad. There seems to be a problem at the edges (it seems to be forcing the reconstruction to be zero at redshift zero and thus changing the shape on the low end), but overall it looks pretty good. (At least much better than it's ever looked before). Thoughts Alexia?


I've chosen the binning which looks the best. Like what happened in the mock data the reconstruction changes a lot depending on the binning. However, it looks like I might actually have something working here! Woo hoo!

Next steps
Higher Statistics: I'm not sure if I should run on Eyal's LRGs or on a bigger set of CAS data. The CAS data is easier to do right away, so I might as well set that going (it will take a couple days), and in the mean time I can work on getting my code to read in Eyal's randoms instead of generating it's own. Will talk to Schlegel about this and see what he thinks.

Other Catalogs: According to Eric Huff, Reiko Nakajima has a SDSS catalog with her own calculated photo-z's that she uses for lensing. I will talk to her about this. I also should get in touch with Antonio Vazquez about catalogs.

Optimization: There is quite a few things I can do to make this code run faster. I need to spend some time writing scripts to simultaneously calculate all the correlation functions (right now they are calculated in succession). This should speed up the code by a factor of ten. I also should write scripts so that I can just set a run going and then log-out of the terminal window.

Organizational Note
For a while now I've been creating a log file for each day with the code I use to generate the plots for the blog post that day. I've been keeping these log files locally on my computer, but decided it might be useful for collaborators to have copies of them, and also for there to be a back-up should something happen to my laptop. I've now added these logs to the repository under: repositoryDir/Jessica/pythonJess/logs. The name of the log file is the date of the blog post which the code corresponds to.

Friday, December 4, 2009

Likelihood Optimatization

I've been playing with various parameters in the likelihood method trying to find the most efficient cuts. The three things I have been playing with are:

added errors (what to use as the input adderr into likelihood_compute)
everything - variability (what happens when we take away variable objects from the L_everything file)
QSO + variability (what happens when we add the variable objects to the L_QSO likelihoods)

Here are my findings (these are on the co-added fluxes, next step is to re-run with single epoch):
Normal Errors (adderr = [0.014, 0.01, 0.01, 0.01, 0.014])
Percent of those targeted are quasars (based on 40/degree^2 targeting)
No variability: 0.649914
Variability everything: 0.656196
Variability everything + QSO = 0.651057

5X Errors (adderr = 5*[0.014, 0.01, 0.01, 0.01, 0.014])
Percent of those targeted are quasars (based on 40/degree^2 targeting)
No variability: 0.641348
Variability everything: 0.645346
Variability everything + QSO = 0.603655

7X Errors (adderr = 7*[0.014, 0.01, 0.01, 0.01, 0.014])
Percent of those targeted are quasars (based on 40/degree^2 targeting)
No variability: 0.627641
Variability everything: 0.624786
Variability everything + QSO = 0.572244

No Errors (adderr = 0.0*[0.014, 0.01, 0.01, 0.01, 0.014])
Percent of those targeted are quasars (based on 40/degree^2 targeting)
No variability: 0.644203
Variability everything: 0.627070
Variability everything + QSO = 0.619075

It looks like the errors we were running have the best numbers, and using the variability everything, but not the variability QSO.

I am going to play more with the definitions of variable everything and variable qso to see if I can get these to work better. I also want to play with not cutting on a L_ratio = 0.01, but perhaps changing L_ratio as a function of L_QSO (it seems we might be able to get a few more objects if we have L_ratio cut decrease as L_QSO gets large.