Showing posts with label PLINK. Show all posts
Showing posts with label PLINK. Show all posts

Wednesday, 6 August 2014

PLINK for Windows

As I'd have to pay for every Gigabyte of internet used, I was a little concerned about working in the comfort of my room. The move to a room with larger window overlooking the backyard where a couple of squirrels and a pair of pigeons reside at the pine tree really helps with a better working at home environment. I'd need to work in the department once in a while, but the alternative is as good as the office, so I feel comforted. And finally, a more sturdy desk chair is provided to me today! Yay!

This is the view from my desk: perfect, isn't it? Getting back to work :)



Back to the topic, since I'm using a laptop installed with Windows 8.1, I was thinking of running PLINK on the laptop, provided I have enough HDD space and RAM, of course. I think with 6GB of RAM, I should be able to do quite a bit before I crash it. A tiny search on PLINK website led me to the MS-DOS version of PLINK, which means I could do the genomic analysis offline without using my group's server! Another yay!

Here's the instruction as per PLINK website on how to download and install PLINK on Windows... I think the links on my page would lead you to the download site, further instructions and manuals for PLINK. Enjoy using the whole genome association analysis toolset!

This page contains some important information regarding how to set up and use PLINK. Individuals familiar with using command line programs can probably skip most of this page.

Download



PLINK
 is now available for free download. Below are links to ZIP files containing binaries compilied on various platforms as well as the C/C++ source code. Linux/Unix users should download the source code and compile (see notes below).


These downloads also contain a version of gPLINK, an (optional) GUI for PLINK. Please see these pages for instructions on use of gPLINK.


Remember This release is considered a stable release, although please remember that we cannot guarantee that it, just like most computer programs, does not contain bugs...


PlatformFileVersion
Linux (x86_64)plink-1.07-x86_64.zipv1.07
Linux (i686)plink-1.07-i686.zipv1.07
MS-DOSplink-1.07-dos.zipv1.07 (to be posted later today, 30-Oct)
Apple Mac (PPC)plink-1.07-mac.zipv1.07 (to be posted next week)
Apple Mac (Intel)plink-1.07-mac-intel.zipv1.07
C/C++ source (.zip)plink-1.07-src.zipv1.07

One more thing... If you download PLINK please either join the very low-volume e-mail list (link from Introduction page) or drop an e-mail to plink AT chgr dot mgh dot harvard dot edu letting me know you've downloaded a copy.


For old versions of PLINK please visit the archive.


Debian users PLINK is available as a Debian package, see these notes. Note, the executable is named snplink in the Debian plink package.

Development version source code



You can download the very latest development source code in this ZIP file. This is really, strongly not recommended for most users. The code posted here could change on a daily basis and is not versioned.
Development source code versions have a p suffix, meaning pre-release. For example, if the current release is 1.04, the next stable release will be 1.05 and the development code will be 1.05p. Note that 1.05 may differ from 1.05p and as noted before, from day-to-day the 1.05 development code may change in any case.
The principle reason for including the source code here is to allow access for specific users to specific, new features. These features are described here.

General installation notes


The PLINK executable file should be placed in either the current working directory or somewhere in the command path. This means that typing
plink


or
./plink


at the command line prompt will run PLINK, no matter which current directory you happen to be in. PLINK is a command line program -- clicking on an icon with the mouse will get you nowhere.
Below, on this page, is a general overview of how to use the command line to run PLINK. The next sections give details about how to install PLINK on different platforms.

Windows/MS-DOS notes



Unzipping the downloaded ZIP file should reveal a single executable program plink.exe. The Windows/MS-DOS version of PLINK is also a command line program, and is run by typing
plink {options...}


not by clicking on the icon with the mouse. Open a DOS windows by selecting "Command Prompt" from the start menu, or entering "command" or "cmd" in the "Run..." option of the start menu.


The folders c:\windows\ or c:\winnt\ are typically in the path, so these are good places to copy the file plink.exe to. You can copy the plink.exe file using Windows, as you would copy-and-paste any file (e.g. using the right-button menu or the keyboard shortcuts control-C (paste) and control-V (paste).


Alternatively, if you know that you will only ever run PLINK on files in a single folder, then you can paste plink.exe into that folder, e.g. C:\work\genetics\. The disadvantage of this approach is that PLINK will not be available from the command line if you are in a folder other than this one.
Once you have copied plink.exe to the correct location, you can test whether or not PLINK is available (i.e. in your command path) by simply typing
plink

at the command line. You should see something like the following message:
     Microsoft Windows XP [Version 5.1.2600]
     (C) Copyright 1985-2001 Microsoft Corp.

     C:\>plink

     @----------------------------------------------------------@
     |         PLINK!       |    v0.99l     |   27/Jul/2006     |
     |----------------------------------------------------------|
     |  (C) 2006 Shaun Purcell, GNU General Public License, v2  |
     |----------------------------------------------------------|
     |       http://pngu.mgh.harvard.edu/purcell/plink/         |
     @----------------------------------------------------------@
 
     Web-based version check ( --noweb to skip )
     Connecting to web...  OK, v0.99l is current
 
     *** Pre-Release Testing Version ***
 
     Writing this text to log file [ plink.log ]
     Analysis started: Fri Jul 28 10:07:57 2006
 
     Options in effect:
 
 
     ERROR: No file [ plink.ped ] exists.

Do not worry about this error message -- normally you would specify your own PED/MAP file names to analyse (i.e. the default input filename is plink.ped).


Please ask your system administrator for help if you do not understand this.


HINT In MS-DOS, you can to increase the width of the window to avoid output lines wrapping around and being hard to read. To do this under Windows XP DOS: right click on the top title/menu bar of the window and select Properties / Layout / Window Size / Width -- increse the width value to a larger value (e.g. 120, or as large as possible without the window getting too big to fit on your screen!).  

Tuesday, 27 May 2014

PLINK: Association Analysis / Accounting for Clusters

It wasn't one of the best days. Had a word with the man above when I bumped into him in the pantry and he didn't seem pleased that it took me two weeks to look at the paper on placing individuals in their geographical location. I'm taking very long because the methods are fundamental, but it wasn't a knowledge that I am born with, unfortunately. I've long suspected that everyone is assumed to be a genius of some sort where I am based and I did feel stupid as I am taking longer time than "most people" to pick up the basics. I know he thinks I'm too slow to learn the skills needed. It's ok, it is not the first time I felt like a complete idiot in this group of geniuses. I have been feeling stupid for a while now. This stupidity is suffocating me from within. I guess this is a price to pay to eventually able to carry the responsibility of having the Cantab post-nominal and the permanent effect of a head damage for ramming into the world of graduate life. Why in the world did I decide to do this? I no longer am in the mood of discern it right now.

Source: http://thegradstudentway.com/blog/wp-content/uploads/2012/07/PhdComics2.jpg

Anyway, back to my learning progress. Any form of progress is better than nothing at all. After the long break from learning, finally I'm back at PLINK tutorial. Today I plotted a multi dimensional scaling (MDS) plot using the HapMap example of two population.

Population stratification is the presence of systematic difference in allele frequencies between subpopulations in a population due to different ancestries, also known as population structure. Stratification analysis use whole genome SNP data to cluster individuals into homogeneous groups. In the tutorial, simple stratification was performed, but the details of it could be found in another chapter in the PLINK website. I think it is worth me spending the whole afternoon going through the main documentation on population stratification as I probably would need to perform this as a routine data treatment procedure and if I don't get it now, I won't get it later. I don't want to form bad habits in programming and jeopardize the quality of my research in future.

If two or more individuals have identical nucleotide sequences in a DNA segment, it is known as identical by state (IBS). If this segment is inherited without recombination by a common ancestor, and is found in two or more individuals, then it is identical by descent (IBD). I find that both IBS and IBD are of the few jargons often mentioned during lab meeting, so it is important that I have these two definition registered in the brain and here. The clustering which I did today was based on pairwise identity-by-state (IBS) distance clustering. No constraint was applied to the process. Usually phenotype criterion and cluster size restriction, plus external matching criteria are specified.

In order to create the MDS plot for the HapMap example of two populations, I first created matrix pairwise IBS distances using this command line:

plink --bfile mydata --cluster --matrix --out myplot

A few files were generated: myplot.mibs, myplot.cluster0, myplot.cluster1, myplot.cluster2, myplot.cluster3, myplot.log. Information are stored in different formation within the four output files resulted from performing the --cluster option.

Using RStudio (it means I used R statistical tool), I created the MDS plot with the code given:

m <- as.matrix (read.table ("myplot.mibs"))
mds <- cmdscale (as.dist (1-m))
k <- c( rep ("purple", 45), rep ("orange", 44) )
plot (mds, pch=20, col=k)

# RStudio was used on Win8 laptop while PLINK was used on UNIX server.

Here's how my plot looked like:

Very interesting combo to use both PLINK and R to generate the plot. According to the tutorial, I could also generate the MDS plot using the --mds-plot option. I have not tried it, so I'm unsure how does it work. I guess it is best that I keep to one which I will be good at, rather than to learn 101 alternatives. I am certain I can beautify my MDS plot. That I can wait.

I better sign off now. I am attending training on "How to Write First Year Report" in the Clinical School later.

Tuesday, 13 May 2014

PLINK by Purcell

Source: http://pngu.mgh.harvard.edu/~purcell/plink/gplink.shtml 

Today's task is simple - to start learning PLINK. The version which I am learning is the command-prompt version of PLINK on UNIX server, which the diagram shows the gPLINK, which is the Java-based software package which can run most of the common PLINK operations. I supposed gPLINK is more user-friendly towards Windows users (like me) who are used to perpetual mouse-clicking rather than using the keyboard. However, it will be more logical for me to utilise our group server, so I'm trying to pick up the UNIX version of it. Right now, I'm learning it using the tutorial provided on the PLINK webpage.

Source: http://pngu.mgh.harvard.edu/~purcell/plink/img/gp_overview2.png
Some brief introduction of PLINK. It was developed by Shaun Purcell, and it's a free open source whole genome association analysis toolset. PLINK itself is used solely to analyse genotype/phenotype data. With the development of gPLINK and Haploview, subsequent visualisation, annotation and storage of results could be performed. By which is a GREAT news for newbies like me!

Let's recount what I find fascinating, which would amuse any normal computer scientist immensely, is that I could call out the command rm to delete the files which I no longer needed on the UNIX server. The ability to finally understand the difference between PED and BED files mentioned on ADMIXTURE. Oh wow! That is a relief! BED is the binary PLINK file which saves space and speeds up subsequent analyses. Tested it using the example dataset "hapmap1":

  1. plink --file hapmap1 --xxxx --xxxx xxxx --out xxxx 
  2. plink --bfile hapmap1 --xxxx --xxxx xxxx --out xxxx

The first one used a normal PLINK file (PED), while the latter used a binary PLINK file (BED). Guess what? The first command took about 5 secs while the second took a sec. It did speed up the analysis! Ok, it is well-known, but it fascinated me.

The first figure is an overview of a structure of the start and end of PLINK really. How PLINK command(s) would eventually generate information which could help others to understand what the scientist has been testing on (or trying to find out). It is important that there is a visualisation of the information, rather than just boring numbers (sorry fellow Mathematicians, I know numbers amuse you, but for general audience, colourful charts still stand out).

Best thing of all, an answer to my previous error, where I need to "apply genotype filter to dataset", appeared after going through the first part of the tutorial. YAY!