spark 有哪些数据类型 https://spark.apache.org/docs/latest/sql-reference.html
Spark 数据类型
Data Types
Spark SQL and DataFrames support the following data types:
- Numeric types
-
ByteType
: Represents 1-byte signed integer numbers. The range of numbers is from-128
to127
. -
ShortType
: Represents 2-byte signed integer numbers. The range of numbers is from-32768
to32767
. -
IntegerType
: Represents 4-byte signed integer numbers. The range of numbers is from-2147483648
to2147483647
. -
LongType
: Represents 8-byte signed integer numbers. The range of numbers is from-9223372036854775808
to9223372036854775807
. -
FloatType
: Represents 4-byte single-precision floating point numbers. -
DoubleType
: Represents 8-byte double-precision floating point numbers. -
DecimalType
: Represents arbitrary-precision signed decimal numbers. Backed internally byjava.math.BigDecimal
. ABigDecimal
consists of an arbitrary precision integer unscaled value and a 32-bit integer scale.
-
- String type
-
StringType
: Represents character string values.
-
- Binary type
-
BinaryType
: Represents byte sequence values.
-
- Boolean type
-
BooleanType
: Represents boolean values.
-
- Datetime type
-
TimestampType
: Represents values comprising values of fields year, month, day, hour, minute, and second. -
DateType
: Represents values comprising values of fields year, month, day.
-
- Complex types
-
ArrayType(elementType, containsNull)
: Represents values comprising a sequence of elements with the type ofelementType
.containsNull
is used to indicate if elements in aArrayType
value can havenull
values. -
MapType(keyType, valueType, valueContainsNull)
: Represents values comprising a set of key-value pairs. The data type of keys are described bykeyType
and the data type of values are described byvalueType
. For aMapType
value, keys are not allowed to havenull
values.valueContainsNull
is used to indicate if values of aMapType
value can havenull
values. -
StructType(fields)
: Represents values with the structure described by a sequence ofStructField
s (fields
).-
StructField(name, dataType, nullable)
: Represents a field in aStructType
. The name of a field is indicated byname
. The data type of a field is indicated bydataType
.nullable
is used to indicate if values of this fields can havenull
values.
-
-
对应的pyspark 数据类型在这里 pyspark.sql.types
一些常见的转化场景:
1. Converts a date/timestamp/string to a value of string, 转成的string 的格式用第二个参数指定
df.withColumn('test', F.date_format(col('Last_Update'),"yyyy/MM/dd")).show()
2. 转成 string后,可以 cast 成你想要的类型,比如下面的 date 型
df = df.withColumn('date', F.date_format(col('Last_Update'),"yyyy-MM-dd").alias('ts').cast("date"))
3. 把 timestamp 秒数(从1970年开始)转成日期格式 string
Ref:
https://*.com/questions/54337991/pyspark-from-unixtime-unix-timestamp-does-not-convert-to-timestamp