Changing default encoding of Python?
I have many "can't encode" and "can't decode" problems with Python when I run my applications from the console. But in the Eclipse PyDev IDE, the default character encoding is set to UTF-8, and I'm fine.
I searched around for setting the default encoding, and people say that Python deletes the
sys.setdefaultencoding function on startup, and we can not use it.
So what's the best solution for it?
Here is a simpler method (hack) that gives you back the
setdefaultencoding() function that was deleted from
import sys # sys.setdefaultencoding() does not exist, here! reload(sys) # Reload does the trick! sys.setdefaultencoding('UTF8')
(Note for Python 3.4+:
reload() is in the
This is not a safe thing to do, though: this is obviously a hack, since
sys.setdefaultencoding() is purposely removed from
sys when Python starts. Reenabling it and changing the default encoding can break code that relies on ASCII being the default (this code can be third-party, which would generally make fixing it impossible or dangerous).
Read more... Read less...
If you get this error when you try to pipe/redirect output of your script
UnicodeEncodeError: 'ascii' codec can't encode characters in position 0-5: ordinal not in range(128)
Just export PYTHONIOENCODING in console and then run your code.
A) To control
python -c 'import sys; print(sys.getdefaultencoding())'
echo "import sys; sys.setdefaultencoding('utf-16-be')" > sitecustomize.py
PYTHONPATH=".:$PYTHONPATH" python -c 'import sys; print(sys.getdefaultencoding())'
You could put your sitecustomize.py higher in your
Also you might like to try
reload(sys).setdefaultencoding by @EOL
B) To control
stdout.encoding you want to set
python -c 'import sys; print(sys.stdin.encoding, sys.stdout.encoding)'
PYTHONIOENCODING="utf-16-be" python -c 'import sys; print(sys.stdin.encoding, sys.stdout.encoding)'
Finally: you can use A) or B) or both!
For earlier versions a solution is to make sure PyDev does not run with UTF-8 as the default encoding. Under Eclipse, run dialog settings ("run configurations", if I remember correctly); you can choose the default encoding on the common tab. Change it to US-ASCII if you want to have these errors 'early' (in other words: in your PyDev environment). Also see an original blog post for this workaround.
Regarding python2 (and python2 only), some of the former answers rely on using the following hack:
import sys reload(sys) # Reload is a hack sys.setdefaultencoding('UTF8')
In my case, it come with a side-effect: I'm using ipython notebooks, and once I run the code the ´print´ function no longer works. I guess there would be solution to it, but still I think using the hack should not be the correct option.
After trying many options, the one that worked for me was using the same code in the
sitecustomize.py, where that piece of code is meant to be. After evaluating that module, the setdefaultencoding function is removed from sys.
So the solution is to append to file
/usr/lib/python2.7/sitecustomize.py the code:
import sys sys.setdefaultencoding('UTF8')
When I use virtualenvwrapper the file I edit is
And when I use with python notebooks and conda, it is
There is an insightful blog post about it.
I paraphrase its content below.
In python 2 which was not as strongly typed regarding the encoding of strings you could perform operations on differently encoded strings, and succeed. E.g. the following would return
u'Toshio' == 'Toshio'
That would hold for every (normal, unprefixed) string that was encoded in
sys.getdefaultencoding(), which defaulted to
ascii, but not others.
The default encoding was meant to be changed system-wide in
site.py, but not somewhere else. The hacks (also presented here) to set it in user modules were just that: hacks, not the solution.
Python 3 did changed the system encoding to default to utf-8 (when LC_CTYPE is unicode-aware), but the fundamental problem was solved with the requirement to explicitly encode "byte"strings whenever they are used with unicode strings.