SlideShare une entreprise Scribd logo
1  sur  30
Scaffold-Based Analytics: Enabling
Hit-to-Lead Decisions by Visualizing
Chemical Series Linked Across
Large Datasets
Deepak Bandyopadhyay,
Constantine Kreatsoulas,
Pat G. Brady, Genaro
Scavello, Dac-Trung Nguyen,
Tyler Peryea, Ajit Jadhav
GSK
NCATS
Thanks to:
Lena Dang and Josh Swamidass (WUSTL),
Rajarshi Guha, Stephen Pickett, Martin
Saunders, Nicola Richmond, Darren Green,
Eric Manas, Todd Graybill, Rob Young, Mike
Ouellette, Stan Martens, Javier Gamo,
Lourdes Rueda
Outline
– Intro: analyzing and merging screening output
– Methods for Scaffold-Based Analytics
– Examples – Linking series across datasets
– Hit Prioritization & Scaffold Hopping (TCAMS)
– Dataset Integration & Scaffold Progression (Kinase “X”)
– Conclusion
2
Small Molecule Lead Discovery at GSK
High Throughput Screening
- Maximize chemical diversity
Focused Screening
- Compound sets tailored
to target families
- Small scale process
Fragment Hit ID
- Low mol weight, ligand
efficient starting points
High-Content / Phenotypic
Screen
- Disease-relevant assays
- Target agnostic
Screening
output: large,
diverse, and
difficult to
navigate
3
GSK,
Tres Cantos,
Spain
DNA Encoded Library
Technology (ELT)
- Massive combinatorial libraries
- Binders found by Next-Gen Seq.
Primary bioassay (pIC50)
Orthogonalassay(pIC50)
Manual Data Surfing
Historical Hit Triage - on Individual Compounds
Criteria
– Activity Data
– Potency in a suite of assays
– Selectivity against off-targets
– Inhibition Frequency Index (IFI)
– Physical/Chemical Properties
– MW, solubility, permeability,…
– Property Forecast Index (PFI)
Use case: isolate good chemical starting points and weed out bad ones
Filters
4
IFI (%) = # HTS assays Hit *100
# HTS assays Tested
PFI = Chromatophic LogD + # of aromatic rings
Lower PFI improves chances of positive outcome
in phys/chem assays correlated with developability
IFI: S. Chakravorty, ACS New Orleans 2013 PFI: R. Young, D.V.S. Green, C. Luscombe, A. Hill. Drug Discovery
Today. Volume 16, Numbers 17/18 September 2011 R
Datasets Used in this Presentation
– Tres Cantos Anti-Malarial Set (TCAMS)
– 13.5k public compounds from GSK HTS
– pIC50 against Plasmodium falciparum (PF)
“susceptible” 3D7 strain
– Percent inhibition against “resistant” DD2 strain
– Other properties including IFI
– In-house data on Kinase “X”
– HTS, FBDD, ELT data
Hit
Prioritization
Dataset
Integration
5
Scaffold
Hopping
?
Outline
– Intro: analyzing and merging screening output
– Methods for Scaffold-Based Analytics
– Examples – Linking series across datasets
– Hit Prioritization & Scaffold Hopping (TCAMS)
– Dataset Integration & Scaffold Progression (Kinase “X”)
– Conclusion
6
Automation is Necessary for Screening Hit Triage…
• Manual selection and scaffold/R-group based SAR do not scale
• 5-50k molecules, 1000’s of chemotypes!
• Traditional methods: clustering, substructure/similarity search, …
SSS2 SSS3SSS1
Manually Merge Results
Multiple Substructure SearchesHierarchical Clustering
Scaffold
Network
(adapted
from J.
Swamidass,
swami.wustl.edu)
7
Agglomerative Clustering
Similarity Search
0.9
0.75
… But Clustering Is Not Sufficient for SAR Navigation
– Agglomerative Clustering:
– Hierarchical Clustering:
– Same underlying issues, adds complexity (level of hierarchy, e.g. # rings)
seals
(fur)
?
singleton
?
ducks
(bill)
?
penguins (flipper)
?
Cluster 3 Cluster 10
similar molecules ≠ same cluster
8
Many singletons
Complete Link Cluster ID
ClusterSize
Molecule  single cluster, can be limiting
Proposed Improvement:
Automatic Decomposition into All (Overlapping) Scaffolds
IFI
1.5%
PF 3D7 LE
0.34
PF 3D7 pIC50
8.1 Molecule
Scaffold(s)
Related Molecules
9
…
49 total
…
226 total
2 total
1.5%
0.318.2
Avg IFI
1.5%
Avg pIC50
8.15
Avg LE
0.32
Avg IFI
3.0%
Avg pIC50
7.8
Avg LE
0.45
Avg IFI
4.0%
Avg pIC50
7.8
Avg LE
0.46
10
Next Step: Combine with Activities and Properties
…
49 total
…
226 total
2 total
1.5%
6.4%
8.5
0.51
0.58
8.2
8.0
2.1%
0.57
7.5
3.0%
0.6
18.1%
24.1%
7.7
0.47
0.36
8.5
2.9%
1.5%
7.4
0.57
0.56
7.9
7.7 8.2
5.0%
0.5
4.4%
0.54
Molecule
Scaffold(s)
Annotation
Related Molecules
– 1
Methods Used to Exhaustively Generate Overlapping
Scaffolds
SSSR scaffolds optimized for R-group tables
Frameworks (GSK) Bemis-Murcko like & RECAP
Exhaustive (pro: complete and con: redundant/too simple)
NCATS
R-Group Tool
4
3
2
Rings
Molecule
Scaffold(s)
Related Molecules
11
Scaffold
Network
Generator
Hierarchical
Directed
Graph of
Scaffolds.
Scales
to large
datasets
Details: Integrating Scaffold-Based Analytics
into a Single Spotfire Visualization
Main Data Table: ChemBLNTD_TCAMS
Compound ID, SMILES, Properties, Activities
Scaffolds from
NCATS R-
Group Tool
Compound
ID
Frames from
Data-Driven
Frameworks
Cluster
from
Clustering
Properties &
activities
aggregated by
scaffold
Framework ID,
FW SMILES,
Cpd IDs
Cluster ID,
Cluster Size,
Cpd IDs
Scaffold info:
IDs, SMILES
Cpd Info: IDs,
SMILES, Properties
Scaffold ID
(many)
Top-Level Scaffold
from Scaffold
Network Generator
scaffold 
subscaffold
Compound
Exemplars from
Top-Level Scaffolds
Scaffold ID
(many)
Scaffold ID
(many)
12
subscaffold
 scaffold
n
n
Method Specific
Group IDs
Molecule
Scaffold(s)
Annotation
Related Molecules
We found
Scaffold
Networks
complex
to integrate
& navigate…
Outline
– Intro: analyzing and merging screening output
– Methods for Scaffold-Based Analytics
– Examples – Linking series across datasets
– Hit Prioritization & Scaffold Hopping (TCAMS)
– Dataset Integration & Scaffold Progression (Kinase “X”)
– Conclusion
13
Framework Overlaps in Related Molecules
Reveal Substructures Associated with Activity
14
Framework
not active in
3D7 strain;
not found by
R-group tool Frameworks
active and
overlapping
Framework
moderately
active
Color by:
Framework
Sector size:
# molecules
Size by:
Ligand
Efficiency
(PF 3D7)
Hit
Prioritization
PercentinhibitioninDD2(PFresistantstrain)
pIC50 in 3D7 (PF susceptible strain)
Each pie is one compound
Each sector/color is one framework
Exemplar compounds
PercentinhibitioninDD2(resistantstrain)
pIC50 in 3D7 (PF susceptible strain)
Scaffold Networks Example: Identify
Related Scaffolds with a Desirable Profile
15
Trellis by:
# rings in
scaffold
Color by:
Top-Level
Scaffold
Size by:
Ligand
Efficiency
(PF 3D7)
Scaffold
Hopping
?
… possibly
more layers
with higher
# rings …
Find new bicyclic and tricyclic scaffolds
active against resistant DD2 strain
Original tricyclic scaffold inactive
against resistant DD2 strain
RINGS = RINGS =
NCATS R-Group Tool Connects Molecules to
Scaffolds with Aggregate Data and Drill-Down
16
– Minimum # of “useful” scaffolds
– Tautomers under single scaffold
Bonus: sensible R-group tables generated
5.7k scaffolds, filtered to 428 by max pIC50
Avg.IFI
Avg. pIC50 in 3D7 (PF sensitive strain)
NCATS R-Group Tool Example:
Deconstruct SAR of Related Molecules
Quinazolines
alone active,
ligand efficient
Discover alt. tricycles
Indazoles
alone only
weakly
active
17
Scaffold
Hopping
?
pIC50 in 3D7 (PF susceptible strain)
IFI
Fuse Design Ideas
Each pie is one compound
Each sector/color is one scaffold
Size by Ligand Efficiency (3D7)
NCATS R-Group Tool Example:
Iterative SAR Exploration
New tricycle scaffold
(1824) seems more
active than indoles or
quinazolines alone
18
pIC50 in 3D7 (PF susceptible strain)
IFI
Scaffold
Hopping
?
Each pie is one compound
Each sector/color is one scaffold
Size by Ligand Efficiency (3D7)
Scaffold-Based Decision Making
and Hit ID Integration
– Kinase “X”
– Candidate compound demonstrates exquisite kinase selectivity
– Active against Wild-Type, Inactive against Mutant enzyme
– Backup program
– New screens analyzed & integrated using NCATS R-Group Tool
19
HTS 2014
350K top-up
3613 pIC50s
HTS 2012
2M screened
4564 pIC50s
2011 2012 2014 (backup)
Fragment
hits
288 pIC50s
DNA ELT
130 libraries
824 features
No activity dataActivity data available
9259
cpds
Goal: identify selective backup series from new Hit ID efforts
Dataset
Integration
HTS 2014 hit
Selective Lead Series Linked Across Datasets
20
MeanΔ(WTpIC50–mutantpIC50)
Mean PFIpred
Scaffold-Level Details:
Mech. pIC50: 7.1
Cell pIC50: 6.3
LE: 0.44
Statistics for 8 exemplars
Mech. pIC50: 6.0 ± 0.88
Cell pIC50: 5.3 ± 0.81
LE: 0.35 ± 0.05
Chemistry initiated on series!
HTS 2012 hit (not followed up)
Scaffold classification by mutant binding
Selective WT/mut.
Non-selective
Size: pIC50
Assay Drill-Down:
Mechanistic
Full-length WT
Truncated WT
Cell
Mutant
pIC50
GSK Compound ID
20122014
Dataset
Integration
Identify and Test Unmeasured Compounds
Based on Overlap with Actives Across Datasets
PFI PFI
MW
Ligand-
efficient
HTS hit
Ligand-efficient
HTS and
fragment hits
21
Dataset
Integration
Weak active for Kinase “X”
Trellis by
Scaffold
Color by LE
Shape by:
Identify and Test Unmeasured Compounds
Based on Overlap with Actives Across Datasets
PFI PFI
MW
Ligand-
efficient
HTS hit
Low
MW/PFI
untested
fragment
Low MW/PFI
ELT feature
to synthesize
Ligand-efficient
HTS and
fragment hits
Low
MW/PFI
untested
fragment
Low MW/PFI
ELT feature
to synthesize
22
Dataset
Integration
Weak active for Kinase “X”
Trellis by
Scaffold
Color by LE
Shape by:
Conclusions and Future Directions
23
• Merging datasets using scaffolds enables a cohesive visualization
of chemical series and suggests opportunities for hybridization
• Automated scaffold and R-group generation is a powerful way to
prioritize hits and replace scaffolds in large and diverse datasets
• Partitioning into clusters is ambiguous, incomplete for SAR navigation.
• Scaffold-Generation Methods (Frameworks, Scaffold Networks,
NCATS R-Group Tool) have their differences, pros and cons
• All methods revealed similar insights from the TCAMS dataset
• Future improvements:
• Scalability to larger and ever-changing datasets
• Automated selection of informative overlapping scaffolds
• Combining multiple scaffold-generation methods
Thank You & Questions
24
Backup and References
– Scaffold Generation Methods:
– NCATS R-group analysis (http://tripod.nih.gov/?p=46 )
– Frameworks (Data-Driven Clustering, GSK/ChemAxon)
– Scaffold Network Generator (http://swami.wustl.edu/sng)
– Agglomerative Clustering (Complete Linkage, GSK/ChemAxon)
25
G. Harper, G. S. Bravi, S. D. Pickett, J. Hussain, and D. V. S.
Green. J. Chem. Inf. Comput. Sci., 44(6), 2145-2156 (2004)
NCATS R–group tool @
http://tripod.nih.gov
M. K. Matlock, J.M. Zaretzki, and S. J. Swamidass.
Bioinformatics. 29(20), 2655-2656 (2013).
Hit Prioritization via Clustering:
Exploration within Pre-determined Groups Only
– ~2000 complete linkage clusters in TCAMS set
– Initial clustering limits neighbors you can discover
Percent inh. in DD2 (PF resistant strain)
IFI
Query molecules (scatter plot)
pXC50 in 3D7 (PF susceptible strain)
#aromaticrings
26
Hit
Prioritization
Using GSK Frameworks
– 80k GSK frameworks, 7.5k RECAP fragments in TCAMS set
– Score of a framework = Average activity of molecules containing it
– Low scoring frameworks can be filtered out
– Issues identified:
– Many equivalent and redundant frameworks
– Tautomers not unified by current implementation
27
Related Molecules with Framework Overlaps:
Reveal Potential Scaffold Hops
Shared framework,
Related chemotypes
Opportunity to design
hybrid series
Color by:
Framework
Sector size:
# molecules
Size by:
Ligand
Efficiency
28
Scaffold
Hopping
?
PercentinhibitioninDD2(PFresistantstrain)
pXC50 in 3D7 (PF susceptible strain)
Molecule
Scaffold(s)
Related Molecules
Each pie is one compound
Each sector/color is one framework
Hit Prioritization via Scaffold Networks:
Navigate to Related Scaffolds
13.5k compounds map to 7715 top-level scaffolds
(28.5k total)
29
Color by:
Top-Level Scaffold
Size by:
Ligand Efficiency
Trellis by:
Number
of rings in
scaffold
Hit
Prioritization
Percent inhibition in DD2 (PF resistant strain)
pXC50in3D7(PFsusceptiblestrain)
2
3
4+
Rings
… possibly more layers with higher # rings …
Related Molecules from NCATS R-Group Tool:
Visualizing Scaffold Overlap and Activity
Co-occurring
active scaffolds
Scaffold 4719
active by itself
Scaffold 978 alone
not highly active
30
pXC50 in 3D7 (PF susceptible strain)
IFI
Hit
Prioritization
Each pie is one compound
Each sector/color is one scaffold

Contenu connexe

Tendances

Update on the Druggable Proteome
Update on the Druggable ProteomeUpdate on the Druggable Proteome
Update on the Druggable ProteomeChris Southan
 
Bioinformatics t9-t10-bio cheminformatics-wimvancriekinge_v2013
Bioinformatics t9-t10-bio cheminformatics-wimvancriekinge_v2013Bioinformatics t9-t10-bio cheminformatics-wimvancriekinge_v2013
Bioinformatics t9-t10-bio cheminformatics-wimvancriekinge_v2013Prof. Wim Van Criekinge
 
Exploring Chemical and Biological Knowledge Spaces with PubChem
Exploring Chemical and Biological Knowledge Spaces with PubChemExploring Chemical and Biological Knowledge Spaces with PubChem
Exploring Chemical and Biological Knowledge Spaces with PubChemPaul Thiessen
 
Cartic Ramakrishnan's dissertation defense
Cartic Ramakrishnan's dissertation defenseCartic Ramakrishnan's dissertation defense
Cartic Ramakrishnan's dissertation defenseCartic Ramakrishnan
 
2015 bioinformatics bio_cheminformatics_wim_vancriekinge
2015 bioinformatics bio_cheminformatics_wim_vancriekinge2015 bioinformatics bio_cheminformatics_wim_vancriekinge
2015 bioinformatics bio_cheminformatics_wim_vancriekingeProf. Wim Van Criekinge
 
Reproducibility in cheminformatics and computational chemistry research: cert...
Reproducibility in cheminformatics and computational chemistry research: cert...Reproducibility in cheminformatics and computational chemistry research: cert...
Reproducibility in cheminformatics and computational chemistry research: cert...Greg Landrum
 
2016 Bio-IT World Cell Line Coordination 2016-04-06v1
2016 Bio-IT World Cell Line Coordination 2016-04-06v12016 Bio-IT World Cell Line Coordination 2016-04-06v1
2016 Bio-IT World Cell Line Coordination 2016-04-06v1Bruce Kozuma
 
Delroy Cameron's Dissertation Defense: A Contenxt-Driven Subgraph Model for L...
Delroy Cameron's Dissertation Defense: A Contenxt-Driven Subgraph Model for L...Delroy Cameron's Dissertation Defense: A Contenxt-Driven Subgraph Model for L...
Delroy Cameron's Dissertation Defense: A Contenxt-Driven Subgraph Model for L...Amit Sheth
 
Analysis with biological pathways:
Analysis with biological pathways: Analysis with biological pathways:
Analysis with biological pathways: Chris Evelo
 
GA4GH Metadata task team presentation
GA4GH Metadata task team presentation GA4GH Metadata task team presentation
GA4GH Metadata task team presentation Melanie Courtot
 
2011-10-11 Open PHACTS at BioIT World Europe
2011-10-11 Open PHACTS at BioIT World Europe2011-10-11 Open PHACTS at BioIT World Europe
2011-10-11 Open PHACTS at BioIT World Europeopen_phacts
 
Towards semantic systems chemical biology
Towards semantic systems chemical biology Towards semantic systems chemical biology
Towards semantic systems chemical biology Bin Chen
 
Opening up pharmacological space, the OPEN PHACTs api
Opening up pharmacological space, the OPEN PHACTs apiOpening up pharmacological space, the OPEN PHACTs api
Opening up pharmacological space, the OPEN PHACTs apiChris Evelo
 
Assessing Drug Safety Using AI
Assessing Drug Safety Using AIAssessing Drug Safety Using AI
Assessing Drug Safety Using AIDatabricks
 
Results Vary: The Pragmatics of Reproducibility and Research Object Frameworks
Results Vary: The Pragmatics of Reproducibility and Research Object FrameworksResults Vary: The Pragmatics of Reproducibility and Research Object Frameworks
Results Vary: The Pragmatics of Reproducibility and Research Object FrameworksCarole Goble
 
Patent chemisty big bang: utilities for SMEs
Patent chemisty big bang: utilities for SMEsPatent chemisty big bang: utilities for SMEs
Patent chemisty big bang: utilities for SMEsChris Southan
 
RARE and FAIR Science: Reproducibility and Research Objects
RARE and FAIR Science: Reproducibility and Research ObjectsRARE and FAIR Science: Reproducibility and Research Objects
RARE and FAIR Science: Reproducibility and Research ObjectsCarole Goble
 

Tendances (20)

Update on the Druggable Proteome
Update on the Druggable ProteomeUpdate on the Druggable Proteome
Update on the Druggable Proteome
 
Bioinformatics t9-t10-bio cheminformatics-wimvancriekinge_v2013
Bioinformatics t9-t10-bio cheminformatics-wimvancriekinge_v2013Bioinformatics t9-t10-bio cheminformatics-wimvancriekinge_v2013
Bioinformatics t9-t10-bio cheminformatics-wimvancriekinge_v2013
 
Exploring Chemical and Biological Knowledge Spaces with PubChem
Exploring Chemical and Biological Knowledge Spaces with PubChemExploring Chemical and Biological Knowledge Spaces with PubChem
Exploring Chemical and Biological Knowledge Spaces with PubChem
 
Cartic Ramakrishnan's dissertation defense
Cartic Ramakrishnan's dissertation defenseCartic Ramakrishnan's dissertation defense
Cartic Ramakrishnan's dissertation defense
 
2015 bioinformatics bio_cheminformatics_wim_vancriekinge
2015 bioinformatics bio_cheminformatics_wim_vancriekinge2015 bioinformatics bio_cheminformatics_wim_vancriekinge
2015 bioinformatics bio_cheminformatics_wim_vancriekinge
 
Reproducibility in cheminformatics and computational chemistry research: cert...
Reproducibility in cheminformatics and computational chemistry research: cert...Reproducibility in cheminformatics and computational chemistry research: cert...
Reproducibility in cheminformatics and computational chemistry research: cert...
 
2016 Bio-IT World Cell Line Coordination 2016-04-06v1
2016 Bio-IT World Cell Line Coordination 2016-04-06v12016 Bio-IT World Cell Line Coordination 2016-04-06v1
2016 Bio-IT World Cell Line Coordination 2016-04-06v1
 
Delroy Cameron's Dissertation Defense: A Contenxt-Driven Subgraph Model for L...
Delroy Cameron's Dissertation Defense: A Contenxt-Driven Subgraph Model for L...Delroy Cameron's Dissertation Defense: A Contenxt-Driven Subgraph Model for L...
Delroy Cameron's Dissertation Defense: A Contenxt-Driven Subgraph Model for L...
 
Analysis with biological pathways:
Analysis with biological pathways: Analysis with biological pathways:
Analysis with biological pathways:
 
GA4GH Metadata task team presentation
GA4GH Metadata task team presentation GA4GH Metadata task team presentation
GA4GH Metadata task team presentation
 
B.3.5
B.3.5B.3.5
B.3.5
 
2011-10-11 Open PHACTS at BioIT World Europe
2011-10-11 Open PHACTS at BioIT World Europe2011-10-11 Open PHACTS at BioIT World Europe
2011-10-11 Open PHACTS at BioIT World Europe
 
Towards semantic systems chemical biology
Towards semantic systems chemical biology Towards semantic systems chemical biology
Towards semantic systems chemical biology
 
Opening up pharmacological space, the OPEN PHACTs api
Opening up pharmacological space, the OPEN PHACTs apiOpening up pharmacological space, the OPEN PHACTs api
Opening up pharmacological space, the OPEN PHACTs api
 
Contrast Pattern Aided Regression and Classification
Contrast Pattern Aided Regression and ClassificationContrast Pattern Aided Regression and Classification
Contrast Pattern Aided Regression and Classification
 
Assessing Drug Safety Using AI
Assessing Drug Safety Using AIAssessing Drug Safety Using AI
Assessing Drug Safety Using AI
 
Results Vary: The Pragmatics of Reproducibility and Research Object Frameworks
Results Vary: The Pragmatics of Reproducibility and Research Object FrameworksResults Vary: The Pragmatics of Reproducibility and Research Object Frameworks
Results Vary: The Pragmatics of Reproducibility and Research Object Frameworks
 
Patent chemisty big bang: utilities for SMEs
Patent chemisty big bang: utilities for SMEsPatent chemisty big bang: utilities for SMEs
Patent chemisty big bang: utilities for SMEs
 
NETTAB 2012
NETTAB 2012NETTAB 2012
NETTAB 2012
 
RARE and FAIR Science: Reproducibility and Research Objects
RARE and FAIR Science: Reproducibility and Research ObjectsRARE and FAIR Science: Reproducibility and Research Objects
RARE and FAIR Science: Reproducibility and Research Objects
 

Similaire à Scaffold-based Analytics: Enabling Hit-to-Lead Decisions by Visualizing Chemical Series Linked Across Large Datasets (ACS Boston 2015)

EnrichNet: Graph-based statistic and web-application for gene/protein set enr...
EnrichNet: Graph-based statistic and web-application for gene/protein set enr...EnrichNet: Graph-based statistic and web-application for gene/protein set enr...
EnrichNet: Graph-based statistic and web-application for gene/protein set enr...Enrico Glaab
 
Integrative analysis of transcriptomics and proteomics data with ArrayMining ...
Integrative analysis of transcriptomics and proteomics data with ArrayMining ...Integrative analysis of transcriptomics and proteomics data with ArrayMining ...
Integrative analysis of transcriptomics and proteomics data with ArrayMining ...Natalio Krasnogor
 
Integrative Networks Centric Bioinformatics
Integrative Networks Centric BioinformaticsIntegrative Networks Centric Bioinformatics
Integrative Networks Centric BioinformaticsNatalio Krasnogor
 
AIQC - ISCB 2022.pdf
AIQC - ISCB 2022.pdfAIQC - ISCB 2022.pdf
AIQC - ISCB 2022.pdfLayne Sadler
 
Population-Based DNA Variant Analysis
Population-Based DNA Variant AnalysisPopulation-Based DNA Variant Analysis
Population-Based DNA Variant AnalysisGolden Helix
 
Omics data integration for MSA | International Society for Clinical Biostatis...
Omics data integration for MSA | International Society for Clinical Biostatis...Omics data integration for MSA | International Society for Clinical Biostatis...
Omics data integration for MSA | International Society for Clinical Biostatis...Said el Bouhaddani 👩‍💻
 
Microarray biotechnologg ppy dna microarrays
Microarray biotechnologg ppy dna microarraysMicroarray biotechnologg ppy dna microarrays
Microarray biotechnologg ppy dna microarraysayeshasattarsandhu
 
Extracting a cellular hierarchy from high-dimensional cytometry data with SPADE
Extracting a cellular hierarchy from high-dimensional cytometry data with SPADEExtracting a cellular hierarchy from high-dimensional cytometry data with SPADE
Extracting a cellular hierarchy from high-dimensional cytometry data with SPADENikolas Pontikos
 
CRISPR Screening: the What, Why and How
CRISPR Screening: the What, Why and HowCRISPR Screening: the What, Why and How
CRISPR Screening: the What, Why and HowHorizonDiscovery
 
A_Pope_RQRM_LeadDisc_June_2016
A_Pope_RQRM_LeadDisc_June_2016A_Pope_RQRM_LeadDisc_June_2016
A_Pope_RQRM_LeadDisc_June_2016Andrew Pope
 
An Overview to Protein bioinformatics
An Overview to Protein bioinformaticsAn Overview to Protein bioinformatics
An Overview to Protein bioinformaticsJoel Ricci-López
 
ppgardner-lecture06-homologysearch.pdf
ppgardner-lecture06-homologysearch.pdfppgardner-lecture06-homologysearch.pdf
ppgardner-lecture06-homologysearch.pdfPaul Gardner
 
scRNA-Seq Workshop Presentation - Stem Cell Network 2018
scRNA-Seq Workshop Presentation - Stem Cell Network 2018scRNA-Seq Workshop Presentation - Stem Cell Network 2018
scRNA-Seq Workshop Presentation - Stem Cell Network 2018David Cook
 
CIBEC Presentation Fatma Sayed.pptx
CIBEC Presentation Fatma Sayed.pptxCIBEC Presentation Fatma Sayed.pptx
CIBEC Presentation Fatma Sayed.pptxFatma Sayed Ibrahim
 
Visual Exploration of Clinical and Genomic Data for Patient Stratification
Visual Exploration of Clinical and Genomic Data for Patient StratificationVisual Exploration of Clinical and Genomic Data for Patient Stratification
Visual Exploration of Clinical and Genomic Data for Patient StratificationNils Gehlenborg
 
Bioinformatics MiRON
Bioinformatics MiRONBioinformatics MiRON
Bioinformatics MiRONPrabin Shakya
 

Similaire à Scaffold-based Analytics: Enabling Hit-to-Lead Decisions by Visualizing Chemical Series Linked Across Large Datasets (ACS Boston 2015) (20)

presentation
presentationpresentation
presentation
 
EnrichNet: Graph-based statistic and web-application for gene/protein set enr...
EnrichNet: Graph-based statistic and web-application for gene/protein set enr...EnrichNet: Graph-based statistic and web-application for gene/protein set enr...
EnrichNet: Graph-based statistic and web-application for gene/protein set enr...
 
Integrative analysis of transcriptomics and proteomics data with ArrayMining ...
Integrative analysis of transcriptomics and proteomics data with ArrayMining ...Integrative analysis of transcriptomics and proteomics data with ArrayMining ...
Integrative analysis of transcriptomics and proteomics data with ArrayMining ...
 
May 15 workshop
May 15  workshopMay 15  workshop
May 15 workshop
 
May workshop
May workshopMay workshop
May workshop
 
Integrative Networks Centric Bioinformatics
Integrative Networks Centric BioinformaticsIntegrative Networks Centric Bioinformatics
Integrative Networks Centric Bioinformatics
 
AIQC - ISCB 2022.pdf
AIQC - ISCB 2022.pdfAIQC - ISCB 2022.pdf
AIQC - ISCB 2022.pdf
 
Population-Based DNA Variant Analysis
Population-Based DNA Variant AnalysisPopulation-Based DNA Variant Analysis
Population-Based DNA Variant Analysis
 
Omics data integration for MSA | International Society for Clinical Biostatis...
Omics data integration for MSA | International Society for Clinical Biostatis...Omics data integration for MSA | International Society for Clinical Biostatis...
Omics data integration for MSA | International Society for Clinical Biostatis...
 
Microarray biotechnologg ppy dna microarrays
Microarray biotechnologg ppy dna microarraysMicroarray biotechnologg ppy dna microarrays
Microarray biotechnologg ppy dna microarrays
 
Extracting a cellular hierarchy from high-dimensional cytometry data with SPADE
Extracting a cellular hierarchy from high-dimensional cytometry data with SPADEExtracting a cellular hierarchy from high-dimensional cytometry data with SPADE
Extracting a cellular hierarchy from high-dimensional cytometry data with SPADE
 
CRISPR Screening: the What, Why and How
CRISPR Screening: the What, Why and HowCRISPR Screening: the What, Why and How
CRISPR Screening: the What, Why and How
 
A_Pope_RQRM_LeadDisc_June_2016
A_Pope_RQRM_LeadDisc_June_2016A_Pope_RQRM_LeadDisc_June_2016
A_Pope_RQRM_LeadDisc_June_2016
 
Practical semantics in the pharmaceutical industry - the Open PHACTS project
Practical semantics in the pharmaceutical industry - the Open PHACTS projectPractical semantics in the pharmaceutical industry - the Open PHACTS project
Practical semantics in the pharmaceutical industry - the Open PHACTS project
 
An Overview to Protein bioinformatics
An Overview to Protein bioinformaticsAn Overview to Protein bioinformatics
An Overview to Protein bioinformatics
 
ppgardner-lecture06-homologysearch.pdf
ppgardner-lecture06-homologysearch.pdfppgardner-lecture06-homologysearch.pdf
ppgardner-lecture06-homologysearch.pdf
 
scRNA-Seq Workshop Presentation - Stem Cell Network 2018
scRNA-Seq Workshop Presentation - Stem Cell Network 2018scRNA-Seq Workshop Presentation - Stem Cell Network 2018
scRNA-Seq Workshop Presentation - Stem Cell Network 2018
 
CIBEC Presentation Fatma Sayed.pptx
CIBEC Presentation Fatma Sayed.pptxCIBEC Presentation Fatma Sayed.pptx
CIBEC Presentation Fatma Sayed.pptx
 
Visual Exploration of Clinical and Genomic Data for Patient Stratification
Visual Exploration of Clinical and Genomic Data for Patient StratificationVisual Exploration of Clinical and Genomic Data for Patient Stratification
Visual Exploration of Clinical and Genomic Data for Patient Stratification
 
Bioinformatics MiRON
Bioinformatics MiRONBioinformatics MiRON
Bioinformatics MiRON
 

Dernier

Machine learning classification ppt.ppt
Machine learning classification  ppt.pptMachine learning classification  ppt.ppt
Machine learning classification ppt.pptamreenkhanum0307
 
Defining Constituents, Data Vizzes and Telling a Data Story
Defining Constituents, Data Vizzes and Telling a Data StoryDefining Constituents, Data Vizzes and Telling a Data Story
Defining Constituents, Data Vizzes and Telling a Data StoryJeremy Anderson
 
Identifying Appropriate Test Statistics Involving Population Mean
Identifying Appropriate Test Statistics Involving Population MeanIdentifying Appropriate Test Statistics Involving Population Mean
Identifying Appropriate Test Statistics Involving Population MeanMYRABACSAFRA2
 
RadioAdProWritingCinderellabyButleri.pdf
RadioAdProWritingCinderellabyButleri.pdfRadioAdProWritingCinderellabyButleri.pdf
RadioAdProWritingCinderellabyButleri.pdfgstagge
 
While-For-loop in python used in college
While-For-loop in python used in collegeWhile-For-loop in python used in college
While-For-loop in python used in collegessuser7a7cd61
 
ASML's Taxonomy Adventure by Daniel Canter
ASML's Taxonomy Adventure by Daniel CanterASML's Taxonomy Adventure by Daniel Canter
ASML's Taxonomy Adventure by Daniel Cantervoginip
 
Student Profile Sample report on improving academic performance by uniting gr...
Student Profile Sample report on improving academic performance by uniting gr...Student Profile Sample report on improving academic performance by uniting gr...
Student Profile Sample report on improving academic performance by uniting gr...Seán Kennedy
 
Multiple time frame trading analysis -brianshannon.pdf
Multiple time frame trading analysis -brianshannon.pdfMultiple time frame trading analysis -brianshannon.pdf
Multiple time frame trading analysis -brianshannon.pdfchwongval
 
GA4 Without Cookies [Measure Camp AMS]
GA4 Without Cookies [Measure Camp AMS]GA4 Without Cookies [Measure Camp AMS]
GA4 Without Cookies [Measure Camp AMS]📊 Markus Baersch
 
NLP Data Science Project Presentation:Predicting Heart Disease with NLP Data ...
NLP Data Science Project Presentation:Predicting Heart Disease with NLP Data ...NLP Data Science Project Presentation:Predicting Heart Disease with NLP Data ...
NLP Data Science Project Presentation:Predicting Heart Disease with NLP Data ...Boston Institute of Analytics
 
Predictive Analysis for Loan Default Presentation : Data Analysis Project PPT
Predictive Analysis for Loan Default  Presentation : Data Analysis Project PPTPredictive Analysis for Loan Default  Presentation : Data Analysis Project PPT
Predictive Analysis for Loan Default Presentation : Data Analysis Project PPTBoston Institute of Analytics
 
Top 5 Best Data Analytics Courses In Queens
Top 5 Best Data Analytics Courses In QueensTop 5 Best Data Analytics Courses In Queens
Top 5 Best Data Analytics Courses In Queensdataanalyticsqueen03
 
INTERNSHIP ON PURBASHA COMPOSITE TEX LTD
INTERNSHIP ON PURBASHA COMPOSITE TEX LTDINTERNSHIP ON PURBASHA COMPOSITE TEX LTD
INTERNSHIP ON PURBASHA COMPOSITE TEX LTDRafezzaman
 
How we prevented account sharing with MFA
How we prevented account sharing with MFAHow we prevented account sharing with MFA
How we prevented account sharing with MFAAndrei Kaleshka
 
科罗拉多大学波尔得分校毕业证学位证成绩单-可办理
科罗拉多大学波尔得分校毕业证学位证成绩单-可办理科罗拉多大学波尔得分校毕业证学位证成绩单-可办理
科罗拉多大学波尔得分校毕业证学位证成绩单-可办理e4aez8ss
 
Student profile product demonstration on grades, ability, well-being and mind...
Student profile product demonstration on grades, ability, well-being and mind...Student profile product demonstration on grades, ability, well-being and mind...
Student profile product demonstration on grades, ability, well-being and mind...Seán Kennedy
 
1:1定制(UQ毕业证)昆士兰大学毕业证成绩单修改留信学历认证原版一模一样
1:1定制(UQ毕业证)昆士兰大学毕业证成绩单修改留信学历认证原版一模一样1:1定制(UQ毕业证)昆士兰大学毕业证成绩单修改留信学历认证原版一模一样
1:1定制(UQ毕业证)昆士兰大学毕业证成绩单修改留信学历认证原版一模一样vhwb25kk
 
Advanced Machine Learning for Business Professionals
Advanced Machine Learning for Business ProfessionalsAdvanced Machine Learning for Business Professionals
Advanced Machine Learning for Business ProfessionalsVICTOR MAESTRE RAMIREZ
 
Heart Disease Classification Report: A Data Analysis Project
Heart Disease Classification Report: A Data Analysis ProjectHeart Disease Classification Report: A Data Analysis Project
Heart Disease Classification Report: A Data Analysis ProjectBoston Institute of Analytics
 
办理学位证纽约大学毕业证(NYU毕业证书)原版一比一
办理学位证纽约大学毕业证(NYU毕业证书)原版一比一办理学位证纽约大学毕业证(NYU毕业证书)原版一比一
办理学位证纽约大学毕业证(NYU毕业证书)原版一比一fhwihughh
 

Dernier (20)

Machine learning classification ppt.ppt
Machine learning classification  ppt.pptMachine learning classification  ppt.ppt
Machine learning classification ppt.ppt
 
Defining Constituents, Data Vizzes and Telling a Data Story
Defining Constituents, Data Vizzes and Telling a Data StoryDefining Constituents, Data Vizzes and Telling a Data Story
Defining Constituents, Data Vizzes and Telling a Data Story
 
Identifying Appropriate Test Statistics Involving Population Mean
Identifying Appropriate Test Statistics Involving Population MeanIdentifying Appropriate Test Statistics Involving Population Mean
Identifying Appropriate Test Statistics Involving Population Mean
 
RadioAdProWritingCinderellabyButleri.pdf
RadioAdProWritingCinderellabyButleri.pdfRadioAdProWritingCinderellabyButleri.pdf
RadioAdProWritingCinderellabyButleri.pdf
 
While-For-loop in python used in college
While-For-loop in python used in collegeWhile-For-loop in python used in college
While-For-loop in python used in college
 
ASML's Taxonomy Adventure by Daniel Canter
ASML's Taxonomy Adventure by Daniel CanterASML's Taxonomy Adventure by Daniel Canter
ASML's Taxonomy Adventure by Daniel Canter
 
Student Profile Sample report on improving academic performance by uniting gr...
Student Profile Sample report on improving academic performance by uniting gr...Student Profile Sample report on improving academic performance by uniting gr...
Student Profile Sample report on improving academic performance by uniting gr...
 
Multiple time frame trading analysis -brianshannon.pdf
Multiple time frame trading analysis -brianshannon.pdfMultiple time frame trading analysis -brianshannon.pdf
Multiple time frame trading analysis -brianshannon.pdf
 
GA4 Without Cookies [Measure Camp AMS]
GA4 Without Cookies [Measure Camp AMS]GA4 Without Cookies [Measure Camp AMS]
GA4 Without Cookies [Measure Camp AMS]
 
NLP Data Science Project Presentation:Predicting Heart Disease with NLP Data ...
NLP Data Science Project Presentation:Predicting Heart Disease with NLP Data ...NLP Data Science Project Presentation:Predicting Heart Disease with NLP Data ...
NLP Data Science Project Presentation:Predicting Heart Disease with NLP Data ...
 
Predictive Analysis for Loan Default Presentation : Data Analysis Project PPT
Predictive Analysis for Loan Default  Presentation : Data Analysis Project PPTPredictive Analysis for Loan Default  Presentation : Data Analysis Project PPT
Predictive Analysis for Loan Default Presentation : Data Analysis Project PPT
 
Top 5 Best Data Analytics Courses In Queens
Top 5 Best Data Analytics Courses In QueensTop 5 Best Data Analytics Courses In Queens
Top 5 Best Data Analytics Courses In Queens
 
INTERNSHIP ON PURBASHA COMPOSITE TEX LTD
INTERNSHIP ON PURBASHA COMPOSITE TEX LTDINTERNSHIP ON PURBASHA COMPOSITE TEX LTD
INTERNSHIP ON PURBASHA COMPOSITE TEX LTD
 
How we prevented account sharing with MFA
How we prevented account sharing with MFAHow we prevented account sharing with MFA
How we prevented account sharing with MFA
 
科罗拉多大学波尔得分校毕业证学位证成绩单-可办理
科罗拉多大学波尔得分校毕业证学位证成绩单-可办理科罗拉多大学波尔得分校毕业证学位证成绩单-可办理
科罗拉多大学波尔得分校毕业证学位证成绩单-可办理
 
Student profile product demonstration on grades, ability, well-being and mind...
Student profile product demonstration on grades, ability, well-being and mind...Student profile product demonstration on grades, ability, well-being and mind...
Student profile product demonstration on grades, ability, well-being and mind...
 
1:1定制(UQ毕业证)昆士兰大学毕业证成绩单修改留信学历认证原版一模一样
1:1定制(UQ毕业证)昆士兰大学毕业证成绩单修改留信学历认证原版一模一样1:1定制(UQ毕业证)昆士兰大学毕业证成绩单修改留信学历认证原版一模一样
1:1定制(UQ毕业证)昆士兰大学毕业证成绩单修改留信学历认证原版一模一样
 
Advanced Machine Learning for Business Professionals
Advanced Machine Learning for Business ProfessionalsAdvanced Machine Learning for Business Professionals
Advanced Machine Learning for Business Professionals
 
Heart Disease Classification Report: A Data Analysis Project
Heart Disease Classification Report: A Data Analysis ProjectHeart Disease Classification Report: A Data Analysis Project
Heart Disease Classification Report: A Data Analysis Project
 
办理学位证纽约大学毕业证(NYU毕业证书)原版一比一
办理学位证纽约大学毕业证(NYU毕业证书)原版一比一办理学位证纽约大学毕业证(NYU毕业证书)原版一比一
办理学位证纽约大学毕业证(NYU毕业证书)原版一比一
 

Scaffold-based Analytics: Enabling Hit-to-Lead Decisions by Visualizing Chemical Series Linked Across Large Datasets (ACS Boston 2015)

  • 1. Scaffold-Based Analytics: Enabling Hit-to-Lead Decisions by Visualizing Chemical Series Linked Across Large Datasets Deepak Bandyopadhyay, Constantine Kreatsoulas, Pat G. Brady, Genaro Scavello, Dac-Trung Nguyen, Tyler Peryea, Ajit Jadhav GSK NCATS Thanks to: Lena Dang and Josh Swamidass (WUSTL), Rajarshi Guha, Stephen Pickett, Martin Saunders, Nicola Richmond, Darren Green, Eric Manas, Todd Graybill, Rob Young, Mike Ouellette, Stan Martens, Javier Gamo, Lourdes Rueda
  • 2. Outline – Intro: analyzing and merging screening output – Methods for Scaffold-Based Analytics – Examples – Linking series across datasets – Hit Prioritization & Scaffold Hopping (TCAMS) – Dataset Integration & Scaffold Progression (Kinase “X”) – Conclusion 2
  • 3. Small Molecule Lead Discovery at GSK High Throughput Screening - Maximize chemical diversity Focused Screening - Compound sets tailored to target families - Small scale process Fragment Hit ID - Low mol weight, ligand efficient starting points High-Content / Phenotypic Screen - Disease-relevant assays - Target agnostic Screening output: large, diverse, and difficult to navigate 3 GSK, Tres Cantos, Spain DNA Encoded Library Technology (ELT) - Massive combinatorial libraries - Binders found by Next-Gen Seq.
  • 4. Primary bioassay (pIC50) Orthogonalassay(pIC50) Manual Data Surfing Historical Hit Triage - on Individual Compounds Criteria – Activity Data – Potency in a suite of assays – Selectivity against off-targets – Inhibition Frequency Index (IFI) – Physical/Chemical Properties – MW, solubility, permeability,… – Property Forecast Index (PFI) Use case: isolate good chemical starting points and weed out bad ones Filters 4 IFI (%) = # HTS assays Hit *100 # HTS assays Tested PFI = Chromatophic LogD + # of aromatic rings Lower PFI improves chances of positive outcome in phys/chem assays correlated with developability IFI: S. Chakravorty, ACS New Orleans 2013 PFI: R. Young, D.V.S. Green, C. Luscombe, A. Hill. Drug Discovery Today. Volume 16, Numbers 17/18 September 2011 R
  • 5. Datasets Used in this Presentation – Tres Cantos Anti-Malarial Set (TCAMS) – 13.5k public compounds from GSK HTS – pIC50 against Plasmodium falciparum (PF) “susceptible” 3D7 strain – Percent inhibition against “resistant” DD2 strain – Other properties including IFI – In-house data on Kinase “X” – HTS, FBDD, ELT data Hit Prioritization Dataset Integration 5 Scaffold Hopping ?
  • 6. Outline – Intro: analyzing and merging screening output – Methods for Scaffold-Based Analytics – Examples – Linking series across datasets – Hit Prioritization & Scaffold Hopping (TCAMS) – Dataset Integration & Scaffold Progression (Kinase “X”) – Conclusion 6
  • 7. Automation is Necessary for Screening Hit Triage… • Manual selection and scaffold/R-group based SAR do not scale • 5-50k molecules, 1000’s of chemotypes! • Traditional methods: clustering, substructure/similarity search, … SSS2 SSS3SSS1 Manually Merge Results Multiple Substructure SearchesHierarchical Clustering Scaffold Network (adapted from J. Swamidass, swami.wustl.edu) 7 Agglomerative Clustering Similarity Search 0.9 0.75
  • 8. … But Clustering Is Not Sufficient for SAR Navigation – Agglomerative Clustering: – Hierarchical Clustering: – Same underlying issues, adds complexity (level of hierarchy, e.g. # rings) seals (fur) ? singleton ? ducks (bill) ? penguins (flipper) ? Cluster 3 Cluster 10 similar molecules ≠ same cluster 8 Many singletons Complete Link Cluster ID ClusterSize Molecule  single cluster, can be limiting
  • 9. Proposed Improvement: Automatic Decomposition into All (Overlapping) Scaffolds IFI 1.5% PF 3D7 LE 0.34 PF 3D7 pIC50 8.1 Molecule Scaffold(s) Related Molecules 9 … 49 total … 226 total 2 total
  • 10. 1.5% 0.318.2 Avg IFI 1.5% Avg pIC50 8.15 Avg LE 0.32 Avg IFI 3.0% Avg pIC50 7.8 Avg LE 0.45 Avg IFI 4.0% Avg pIC50 7.8 Avg LE 0.46 10 Next Step: Combine with Activities and Properties … 49 total … 226 total 2 total 1.5% 6.4% 8.5 0.51 0.58 8.2 8.0 2.1% 0.57 7.5 3.0% 0.6 18.1% 24.1% 7.7 0.47 0.36 8.5 2.9% 1.5% 7.4 0.57 0.56 7.9 7.7 8.2 5.0% 0.5 4.4% 0.54 Molecule Scaffold(s) Annotation Related Molecules
  • 11. – 1 Methods Used to Exhaustively Generate Overlapping Scaffolds SSSR scaffolds optimized for R-group tables Frameworks (GSK) Bemis-Murcko like & RECAP Exhaustive (pro: complete and con: redundant/too simple) NCATS R-Group Tool 4 3 2 Rings Molecule Scaffold(s) Related Molecules 11 Scaffold Network Generator Hierarchical Directed Graph of Scaffolds. Scales to large datasets
  • 12. Details: Integrating Scaffold-Based Analytics into a Single Spotfire Visualization Main Data Table: ChemBLNTD_TCAMS Compound ID, SMILES, Properties, Activities Scaffolds from NCATS R- Group Tool Compound ID Frames from Data-Driven Frameworks Cluster from Clustering Properties & activities aggregated by scaffold Framework ID, FW SMILES, Cpd IDs Cluster ID, Cluster Size, Cpd IDs Scaffold info: IDs, SMILES Cpd Info: IDs, SMILES, Properties Scaffold ID (many) Top-Level Scaffold from Scaffold Network Generator scaffold  subscaffold Compound Exemplars from Top-Level Scaffolds Scaffold ID (many) Scaffold ID (many) 12 subscaffold  scaffold n n Method Specific Group IDs Molecule Scaffold(s) Annotation Related Molecules We found Scaffold Networks complex to integrate & navigate…
  • 13. Outline – Intro: analyzing and merging screening output – Methods for Scaffold-Based Analytics – Examples – Linking series across datasets – Hit Prioritization & Scaffold Hopping (TCAMS) – Dataset Integration & Scaffold Progression (Kinase “X”) – Conclusion 13
  • 14. Framework Overlaps in Related Molecules Reveal Substructures Associated with Activity 14 Framework not active in 3D7 strain; not found by R-group tool Frameworks active and overlapping Framework moderately active Color by: Framework Sector size: # molecules Size by: Ligand Efficiency (PF 3D7) Hit Prioritization PercentinhibitioninDD2(PFresistantstrain) pIC50 in 3D7 (PF susceptible strain) Each pie is one compound Each sector/color is one framework Exemplar compounds
  • 15. PercentinhibitioninDD2(resistantstrain) pIC50 in 3D7 (PF susceptible strain) Scaffold Networks Example: Identify Related Scaffolds with a Desirable Profile 15 Trellis by: # rings in scaffold Color by: Top-Level Scaffold Size by: Ligand Efficiency (PF 3D7) Scaffold Hopping ? … possibly more layers with higher # rings … Find new bicyclic and tricyclic scaffolds active against resistant DD2 strain Original tricyclic scaffold inactive against resistant DD2 strain RINGS = RINGS =
  • 16. NCATS R-Group Tool Connects Molecules to Scaffolds with Aggregate Data and Drill-Down 16 – Minimum # of “useful” scaffolds – Tautomers under single scaffold Bonus: sensible R-group tables generated 5.7k scaffolds, filtered to 428 by max pIC50 Avg.IFI Avg. pIC50 in 3D7 (PF sensitive strain)
  • 17. NCATS R-Group Tool Example: Deconstruct SAR of Related Molecules Quinazolines alone active, ligand efficient Discover alt. tricycles Indazoles alone only weakly active 17 Scaffold Hopping ? pIC50 in 3D7 (PF susceptible strain) IFI Fuse Design Ideas Each pie is one compound Each sector/color is one scaffold Size by Ligand Efficiency (3D7)
  • 18. NCATS R-Group Tool Example: Iterative SAR Exploration New tricycle scaffold (1824) seems more active than indoles or quinazolines alone 18 pIC50 in 3D7 (PF susceptible strain) IFI Scaffold Hopping ? Each pie is one compound Each sector/color is one scaffold Size by Ligand Efficiency (3D7)
  • 19. Scaffold-Based Decision Making and Hit ID Integration – Kinase “X” – Candidate compound demonstrates exquisite kinase selectivity – Active against Wild-Type, Inactive against Mutant enzyme – Backup program – New screens analyzed & integrated using NCATS R-Group Tool 19 HTS 2014 350K top-up 3613 pIC50s HTS 2012 2M screened 4564 pIC50s 2011 2012 2014 (backup) Fragment hits 288 pIC50s DNA ELT 130 libraries 824 features No activity dataActivity data available 9259 cpds Goal: identify selective backup series from new Hit ID efforts Dataset Integration
  • 20. HTS 2014 hit Selective Lead Series Linked Across Datasets 20 MeanΔ(WTpIC50–mutantpIC50) Mean PFIpred Scaffold-Level Details: Mech. pIC50: 7.1 Cell pIC50: 6.3 LE: 0.44 Statistics for 8 exemplars Mech. pIC50: 6.0 ± 0.88 Cell pIC50: 5.3 ± 0.81 LE: 0.35 ± 0.05 Chemistry initiated on series! HTS 2012 hit (not followed up) Scaffold classification by mutant binding Selective WT/mut. Non-selective Size: pIC50 Assay Drill-Down: Mechanistic Full-length WT Truncated WT Cell Mutant pIC50 GSK Compound ID 20122014 Dataset Integration
  • 21. Identify and Test Unmeasured Compounds Based on Overlap with Actives Across Datasets PFI PFI MW Ligand- efficient HTS hit Ligand-efficient HTS and fragment hits 21 Dataset Integration Weak active for Kinase “X” Trellis by Scaffold Color by LE Shape by:
  • 22. Identify and Test Unmeasured Compounds Based on Overlap with Actives Across Datasets PFI PFI MW Ligand- efficient HTS hit Low MW/PFI untested fragment Low MW/PFI ELT feature to synthesize Ligand-efficient HTS and fragment hits Low MW/PFI untested fragment Low MW/PFI ELT feature to synthesize 22 Dataset Integration Weak active for Kinase “X” Trellis by Scaffold Color by LE Shape by:
  • 23. Conclusions and Future Directions 23 • Merging datasets using scaffolds enables a cohesive visualization of chemical series and suggests opportunities for hybridization • Automated scaffold and R-group generation is a powerful way to prioritize hits and replace scaffolds in large and diverse datasets • Partitioning into clusters is ambiguous, incomplete for SAR navigation. • Scaffold-Generation Methods (Frameworks, Scaffold Networks, NCATS R-Group Tool) have their differences, pros and cons • All methods revealed similar insights from the TCAMS dataset • Future improvements: • Scalability to larger and ever-changing datasets • Automated selection of informative overlapping scaffolds • Combining multiple scaffold-generation methods
  • 24. Thank You & Questions 24
  • 25. Backup and References – Scaffold Generation Methods: – NCATS R-group analysis (http://tripod.nih.gov/?p=46 ) – Frameworks (Data-Driven Clustering, GSK/ChemAxon) – Scaffold Network Generator (http://swami.wustl.edu/sng) – Agglomerative Clustering (Complete Linkage, GSK/ChemAxon) 25 G. Harper, G. S. Bravi, S. D. Pickett, J. Hussain, and D. V. S. Green. J. Chem. Inf. Comput. Sci., 44(6), 2145-2156 (2004) NCATS R–group tool @ http://tripod.nih.gov M. K. Matlock, J.M. Zaretzki, and S. J. Swamidass. Bioinformatics. 29(20), 2655-2656 (2013).
  • 26. Hit Prioritization via Clustering: Exploration within Pre-determined Groups Only – ~2000 complete linkage clusters in TCAMS set – Initial clustering limits neighbors you can discover Percent inh. in DD2 (PF resistant strain) IFI Query molecules (scatter plot) pXC50 in 3D7 (PF susceptible strain) #aromaticrings 26 Hit Prioritization
  • 27. Using GSK Frameworks – 80k GSK frameworks, 7.5k RECAP fragments in TCAMS set – Score of a framework = Average activity of molecules containing it – Low scoring frameworks can be filtered out – Issues identified: – Many equivalent and redundant frameworks – Tautomers not unified by current implementation 27
  • 28. Related Molecules with Framework Overlaps: Reveal Potential Scaffold Hops Shared framework, Related chemotypes Opportunity to design hybrid series Color by: Framework Sector size: # molecules Size by: Ligand Efficiency 28 Scaffold Hopping ? PercentinhibitioninDD2(PFresistantstrain) pXC50 in 3D7 (PF susceptible strain) Molecule Scaffold(s) Related Molecules Each pie is one compound Each sector/color is one framework
  • 29. Hit Prioritization via Scaffold Networks: Navigate to Related Scaffolds 13.5k compounds map to 7715 top-level scaffolds (28.5k total) 29 Color by: Top-Level Scaffold Size by: Ligand Efficiency Trellis by: Number of rings in scaffold Hit Prioritization Percent inhibition in DD2 (PF resistant strain) pXC50in3D7(PFsusceptiblestrain) 2 3 4+ Rings … possibly more layers with higher # rings …
  • 30. Related Molecules from NCATS R-Group Tool: Visualizing Scaffold Overlap and Activity Co-occurring active scaffolds Scaffold 4719 active by itself Scaffold 978 alone not highly active 30 pXC50 in 3D7 (PF susceptible strain) IFI Hit Prioritization Each pie is one compound Each sector/color is one scaffold

Notes de l'éditeur

  1. Data visualization & exploration environment (we use Spotfire). PFI lipo akin to cLogP. Lower is better. 30 sec.
  2. Adding the hier does not fix the agg isues, only adds complexity in navigation . Things at different levels may not be matched
  3. What I will be describing is a method that exhaustively finds all possible shared (or common or frequent) substructures – which we call scaffolds within your data set using a tool from the NIH. Here is a screening hit that I will use to demonstrate this. … (don’t need to go into gory details) Biaryl substructure is contained in these molecules that have low similarity to the original hit molecule.
  4. We can aggregate activities & properties at the scaffold-level and then drill-down to the underlying data for individual compounds to progress scaffolds of interest.
  5. Text up top. Grey out clustering. Purple box for aggregate props.
  6. Preprocessed substructure search: which substructure encodes activity?
  7. 10 sec. short script
  8. We used the scaffolds to merge all of this data and identify more series that bind selectively
  9. Key message: prioritize ELT with no activity data, just based on overlap with actives from other datasets
  10. This slide can be backup.
  11. Automated substructure search to find part of molecule that’s active. Backup?
  12. Backup
  13. Backup?