Posts Character-Set과 Encoding
Post
Cancel

Character-Set과 Encoding

Charset VS Encoding

  • Before the invention of Unicode, there was no real difference between charset and encoding.
  • Both terms were used to refer to a way to represent letters in binary.
  • Therefore, ASCII, Latin1, Cp1252 etc. can be considered as character sets and encodings at the same time, hence the confusion.

  • 유니코드가 발명되기 전에는 문자 세트와 인코딩 사이에 실질적인 차이가 없었습니다.
  • 두 용어 모두 문자를 이진수로 표현하는 방식을 가리키는 데 사용되었습니다.
  • 따라서 ASCII, Latin1, Cp1252 등은 문자 집합인 동시에 인코딩으로 간주하여 혼동을 일으킬 수 있습니다.
  • 유니코드가 등장하자 문자를 바이트로 표현하는 것이 이제 2단계 프로세스가 되었기 때문에 명확한 구별이 이루어졌습니다.
    • 문자/문자 개념을 코드 포인트(code point)라는 숫자에 연관시킵니다.
    • 비트를 사용하여 이 “코드 포인트” 번호를 인코딩합니다.

따라서 오늘날 문자 집합, 인코딩 두 용어를 다음과 같이 정의할 수 있습니다.

  • 문자 집합
    • 코드 포인트라고 하는 이론적이고 추상적인 숫자에 매핑된 문자 세트입니다.
    • 유니코드는 전 세계에서 사용되는 거의 모든 문자를 포함하는 문자 집합의 예입니다.
    • Charset = letter -> code point
  • 인코딩 체계
    • 이러한 코드 포인트가 UTF-8 또는 UTF-16BE와 같이 바이트로 표현되는 방식을 설명합니다.
    • Encoding scheme = code point -> bytes
  • Here is an example with the euro symbol “€”: image

참고 자료

  • https://medium.com/@joffrey.bion/charset-encoding-encryption-same-thing-6242c3f9da0c
  • https://www.joelonsoftware.com/2003/10/08/the-absolute-minimum-every-software-developer-absolutely-positively-must-know-about-unicode-and-character-sets-no-excuses/
This post is licensed under CC BY 4.0 by the author.

Maven 맛보기

MySQL에서 4bytes 이모지가 물음표로 저장되는 이슈