• Cb: New Blogs #13

    The Cb software is still holding… I jettinsoned the old post cache, which speeded up the processing of blogs considerably, but the system just doesn’t scale right. Yet, Euan has done a great job, and the Cb site has now been online for some three years! Here are some new blogs included in the aggregation and analysis:

  • Aggregating my blogs on this blog using labels

    I am reorganizing my blogs, and will now take advantage of the labels concept of Blogger. The new URLs are:

  • Program for the RDF symposium at the American Chemical Society fall meeting

    It is my great pleasure to present the full symposium program for the RDF session at the American Chemical Society at the Boston meeting in August:

  • Critical mass for Open Notebook Science wikis by prepopulation with RDF data

    One of the problems with Open Data, - Source, and Standards (ODOSOS) is to get critical mass: for the project to move forward, you need both a good user community and an active developers community. The first is the crucial reward for the latter: people using your work is the incentive. Sometimes, a project does not succeed in doing that; Rich was questioning if his ChemPedia project built up enough mass.

  • All CrystalEye data available as PDDL

    I do not think I have seen this earlier, but I do not visit the homepage regularly. But congratulations to the CrystalEye team for releasing the data explicitly as PDDL! I cannot stress enough how useful it is if you add a statement like this to your public domain data. Well done!

  • Looking at your statistical models...

    I do not think I have ever blogged the paper that played an important role in my thesis (doi:10.1021/ci990038z); research of one of the papers in my thesis, started with the hypothesis proposed therein. The paper had a really good idea; but, unfortunately, it did not contain the data to support the hypothesis. That gets me to one important lesson I learned: a QSAR data set of less than 100 molecules is not enough to make untargeted statistical models.

  • QSPR modeling with signatures

    I had to dig deep to find posts on QSAR modeling. There are quite a few on QSAR in Bioclipse, but that focuses on the descriptor calculation. In a quick scan, I could only spot two modeling posts: