显示标签为“Python”的博文。显示所有博文
显示标签为“Python”的博文。显示所有博文

2014年3月4日星期二

Setup Python to work with MS SQL Server in CentOS 6

I spent quite sometime and today finally I was able to use Python scripts to connect to database and run queries using pandas. Here are the steps:

BTW, these instructions are for CentOS 6.4, Python-2.7.6 with pip already installed.

1. Install the required packages.

yum install gcc gcc-c++ python-devel freetds unixODBC unixODBC-devel

Then, you can go ahead and install pyodbc with:

pip install pyodbc

2. Configure FreeTDS

Edit the file /etc/freetds.conf or ~/.freetds.conf if you do not have root privilege.

For each database you will be working with, add the section:

[DB_SERVER]
host = URL
port = 1433
tds version = 7.0

Then, copy the configure file to your home folder:
cp /etc/freetds.conf ~/.freetds.conf

To see if everything is working as of now, try:
tsql -S DB_SERVER -U username -P password

If you can successfully log in and run some simple queries, you know FreeTDS is working properly.

3. Configure unixODBC to work with FreeTDS

Add the following section to /etc/odbcinst.ini:

[FreeTDS]
Description     = MS SQL database access with Free TDS
Driver          = /usr/local/lib/libtdsodbc.so
Setup           = /usr/lib64/libtdsS.so
CPTimeout = 
CPReuse = 
FileUsage = 1
The path might be different for different machines.

Then, for each database you want to work with, add the following section to /etc/odbc.ini or ~/.odbc.ini if you do not have root privilege:

[DB_SOURCE]
Driver = FreeTDS
Description = ODBC connection via FreeTDS
Trace = No
Servername = DB_SERVER
Database = DB_NAME

To see if the configuration is working, try:
isql -v DB_SOURCE  username password

If you can successfully log in and run some simple queries, you know unixODBC is working properly.


Note: to make sure ODBC look for your local config file, do this:
export ODBCINI=/HOMEDIR/.odbc.ini

4. Try to connect with Python

Finally, in Python, your script should look like this:

import pyodbc

dsn = 'DB_SOURCE'
user = 'username'
password = 'password'
database = 'DB_NAME'

con_string = 'DSN=%s;UID=%s;PWD=%s;DATABASE=%s;' % (dsn, user, password, database)
cnxn = pyodbc.connect(con_string)

Now you can embed your SQL queries in Python XD

2014年1月24日星期五

Setting up Python-2.7 on CentOS 6

I just got a Linux desktop with CentOS 6 installed, and I need to set up the environment as Python-2.7 + Pandas + Numpy + Scipy with Emacs 24.3

It turns out that it's very tricky to use yum to install ANYTHING... (at least compared to Ubuntu)

First of all, install Python-2.7 side by side with the original 2.6 since otherwise you will mess up your OS. Then, install pip and configure it to Python-2.7. The step by step instruction is here:
https://github.com/0xdata/h2o/wiki/Installing-python-2.7-on-centos-6.3.-Follow-this-sequence-exactly-for-centos-machine-only

http://toomuchdata.com/2012/06/25/how-to-install-python-2-7-3-on-centos-6-2/

It's important to run:
yum groupinstall "Development tools"
yum install zlib-devel bzip2-devel openssl-devel ncurses-devel sqlite-devel readline-devel tk-devel
Before compiling and installing Python.

Then, by typing:

pip install PACKAGE_NAME

You will be able to install the package for Python-2.7.

For Scipy, the important package to install before it are:
yum install blas blas-devel lapack lapack-devel atlas atlas-devel

For Matpoltlib, the important package to install before it are:
yum install freetype-devel libpng-devel

For iPython to have [Tab] nationalities:
pip install readline

To install Emacs 24:

http://vitalvastness.wordpress.com/2013/07/03/installing-emacs-24-on-centos-6/comment-page-1/

Install liblockfile from here (that’s the x86_64 link) … if you click on the download link it will invoke the package manager and install directly from Firefox.
cd /etc/yum.repos.d
yum install emacs-24.2-4.el6.x86_64

2009年3月18日星期三

Python Looping小技巧

Python下面的循环有很多看起来花哨其实很实用的技巧。首当其冲的就是enumerate(),花哨到用语言很难讲清楚...小例子如下:

for i, v in enumerate(['tic', 'tac', 'toe']):
print i, v

0 tic
1 tac
2 toe

在编程的时候大家经常要用到的情况,就是控制循环的index同时是用来访问某一个数组的。比如用Matlab实现:

arr = {'tic', 'tac', 'toe'};

for i = 1:length(arr)
fprintf('%d %s', i, arr{i});
end

相比之下,大家就知道我为什么说enumerate()实在是很fancy的一个东西。

但enumerate()和我们一直理解的循环有一个很重要的区别,那就是循环的进行不是靠i的值控制的,而是靠enumerate来控制的。咱们写循环的时候一个常见的错误就是循环套循环的时候经常用一样的变量“i”来控制,这样就会出很诡异很诡异的问题,debug都不容易找到。但是在Python下面:

for i, season in enumerate(['Spring', 'Summer', 'Fall', 'Winter']):
print i, season
for i, day in enumerate(['Wed', 'Thu', 'Fri', 'Sat']):
print i, day

0 Spring
0 Wed
1 Thu
2 Fri
3 Sat
1 Summer
0 Wed
1 Thu
2 Fri
3 Sat
2 Fall
0 Wed
1 Thu
2 Fri
3 Sat
3 Winter
0 Wed
1 Thu
2 Fri
3 Sat

虽然我在两层循环用了一样的i,但Python根本不鸟这个,只管用enumerate()生成(index, arr[index])对,生成完之后循环结束。如果我用matlab这样写:

S = {'Spring', 'Summer', 'Fall', 'Winter'};
D = {'Wed', 'Thu', 'Fri', 'Sat'};

for i = 1:length(S)
disp(S{i});
for i = 1:length(D)
disp(D{i});
end
end

大家都知道什么后果了吧...

Enumerate还有一个非常Fancy的例子:

fp = open("file", "r")
for i, line in enumerate(fp):
print "line number: " + lineno + ": " + line.rstrip()

可以一行一行把一个文件读出来,而且很省内存,只要你的内存能装下那一行就行了。在需要把一个超大文本文件里的数据读出来处理的时候,这个例子非常有用。

最后再帖一个enumerate之外控制循环的例子,也是相当fancy的:

questions = ['name', 'quest', 'favorite color']
answers = ['lancelot', 'the holy grail', 'blue']
for q, a in zip(questions, answers):
print 'What is your ' + q + '?', 'It is ' + a + '.'

What is your name? It is lancelot.
What is your quest? It is the holy grail.
What is your favorite color? It is blue.

2008年6月3日星期二

Python变量存储

这次毕设虽然题目挺无聊的,但是为了抓住这个机会练习一下包括Python在内的很多东西,我多少还是认真完成的。因为我的毕设用的是“过完备”的特征提取,所以一张图片的特征维数达到了412160之多...这样所有训练图片的特征一次性存在内存里是不可能的,所以我就使用了各种各样的数据存取操作。我现在了解到的Python相比Matlab唯一的一点美中不足就是数据结构不够统一所导致的数据存储没有统一的函数。从最最基本的以ASCII格式存储矩阵的io.save,到man里面很是推崇的cPickle。我这次各种需求的数据都尝试过了,所以小小的总结一段:

1、对于一个或多个数据类型不同且体积较小(比如1000维以内),而且含有Python特有的数据类型比如list,字典等等的变量,统一变成一个字典然后使用cPickle.dump。这样将来load出来数据类型不变,会节约大量的格式转换代码。

2、对于较大的变量,比如我使用的图像特征,双精度1*412160,以及训练时使用的一个核心矩阵,int8整数,375*412160,无论使用ASCII还是cPickle都会速度非常慢,且极其消耗硬盘空间。这时就要考虑是用2进制来存储。我之前一直纠缠于这个问题甚至最后自己写了一个用ASCII读取int8的函数,就是因为我在使用scipy.io.savemat存储int8类的矩阵时,读出来仍然是双精度,那可是375*412160的矩阵啊,我的内存当时就崩了...后来查了一下scipy.io.mio的源代码,原来里面有这样一段:
np_to_mtypes = {
'f8': miDOUBLE,
'c32': miDOUBLE,
'c24': miDOUBLE,
'c16': miDOUBLE,
'f4': miSINGLE,
'c8': miSINGLE,
'i4': miINT32,
'i2': miINT16,
'u2': miUINT16,
'u1': miUINT8,
'S1': miUINT8,
}
这相当于python到matlab数据类型的一个转换码表。我注意到'u1': miUINT8这一段,意识到savemat是可以保留一些数据类型的。
所以应该:
a = ones(5, dtype = 'u1')
scipy.io.savemat('test.mat', {'a':a})
这样load出来就完全没有问题了...只是要注意,uint8是0~255的。
所以之前一直有问题的原因就是python下面的int8在matlab里是不存在的,所以mio会自动转化为双精度来存储。但是我又必须使用int8,因为int8是-128~+127的。不过还好Python的格式转换不会占用多余的内存。
所以,对于较大型的数据矩阵,使用scipy.io.mio,存取速度非常快,有压缩,节约硬盘空间。注意的话,可妥善保存数据类型。

使用Python PIL的show()显示图片

Python的标准图像库里有一个show()函数,总是不能用。因为他调用了xv,但xv在后面的ubuntu版本中xv都不装了。
sudo ln -s /usr/bin/display /usr/bin/xv
先装一个ImageMagic,在这样一下,就相当于把xv的入口换成了ImageMagic的display。